jobs Logo
Kubex logo

Staff Software Engineer - Kubernetes Observability

Kubex3 days ago
Canada
Senior Level
Full-Time

Top Benefits

Equity
Competitive compensation
Benefits

About the role

About Kubex

Kubex is building the future of autonomous, AI-driven infrastructure optimization. Our platform enables intelligent, policy-driven optimization across Kubernetes, cloud, and GPU-backed environments improving performance, reducing cost, and eliminating waste for some of the world’s most sophisticated technology organizations.

As AI workloads increasingly run on Kubernetes, especially for inference at scale, Kubex is expanding its GPU support to provide more advanced optimization and automation aligned with the unique challenges of GPU-accelerated infrastructure. We combine deep systems expertise, advanced analytics, and patented optimization technology to help customers run AI workloads efficiently and reliably in real-world production environments.

Role Overview Kubex is seeking a Technical Lead, Kubernetes Observability to own the architecture, technical direction and development of our Kubernetes observability capabilities. This is a senior technical leadership role for someone who can define how telemetry is collected, enriched, processed, and used to support infrastructure optimization across complex customer environments.

You will lead the design and evolution of systems that capture workload behavior, resource usage, performance, and operational context across Kubernetes clusters. This includes making key architectural decisions, establishing engineering patterns, prototyping new approaches, and contributing directly to the most critical and technically challenging parts of the implementation.

This role combines high-level ownership with strong hands-on engineering. You will use modern AI-assisted development workflows to accelerate implementation, testing, investigation, and iteration, while remaining accountable for system design, technical decisions, code quality, and production reliability.

The ideal candidate has deep experience building observability or telemetry systems for Kubernetes and is comfortable working across metrics, events, logs, traces, workload metadata, and distributed data collection. Experience with GPU or AI workloads is valuable but not required. You will help build the observability foundation that supports Kubex’s current Kubernetes optimization capabilities and its continued expansion into GPU and AI infrastructure.

Key Responsibilities Lead the design of systems that collect telemetry and improve performance of inference workloads. Contribute directly to production code, remaining deeply hands-on in the design, implementation, and evolution of core platform components. Collaborate closely with other senior engineers, product managers and engineering leadership to coordinate and execute complex software development initiatives. Prototype, validate, and productionize new technical approaches related to AI workload observability and performance optimization. Identify opportunities to extend Kubex’s value beyond inference workloads, including potential future optimizations for training or hybrid workloads.

Required Qualifications 7+ years of professional software engineering experience, including significant experience building production software on Kubernetes. Strong experience designing or building observability and telemetry solutions using technologies such as Prometheus, OpenTelemetry, or similar platforms. Deep understanding of Kubernetes workloads, resource management, API interactions, and the operational challenges of running software across diverse customer clusters. Strong coding skills, preferably in Go, with experience building scalable, reliable, and testable distributed systems. Demonstrated ability to own technical architecture while remaining hands-on with implementation, prototyping, debugging, and production delivery.

Preferred Qualifications Experience building collectors, exporters, agents, or telemetry-processing pipelines for Kubernetes environments. Knowledge of performance analysis, resource optimization, scheduling, or infrastructure efficiency. Exposure to GPU-backed infrastructure, AI inference workloads, or GPU observability technologies.

Why Join Kubex? Play a key role in shaping the future of AI infrastructure optimization. Work on technically challenging problems at the intersection of Kubernetes, GPUs, and AI workloads. Collaborate with a highly experienced, deeply technical team. Influence product direction, architecture, and external technical positioning. Flexible, remote-first culture focused on impact and innovation. Competitive compensation, equity, and benefits.

About Kubex

Software Development
51-200 employees
Founded in 2022

Kubex uses AI-driven automation to continuously optimize Kubernetes, AI/GPU workloads, and cloud infrastructure. By eliminating manual resource tuning and safely adapting to real workload behavior, Kubex reduces cost, improves efficiency, and maintains application stability across complex, dynamic environments.

Similar Jobs