Job Description
Join the 2026 Engineering Revolution
Nebula Systems is at the forefront of the next technological paradigm. We are seeking a visionary Senior AI Infrastructure Engineer to build the scalable backbone of our neural networks for the year 2026 and beyond. If you are passionate about optimizing large-scale distributed systems and preparing the architecture for the future of Artificial General Intelligence, we want to hear from you.
In this role, you will bridge the gap between cutting-edge AI research and robust, production-grade engineering. You will be responsible for architecting high-throughput, low-latency pipelines that power our next-generation applications.
Responsibilities
- Design and maintain high-throughput, low-latency inference pipelines for large language models.
- Optimize GPU clusters and heterogeneous computing environments for distributed training and inference.
- Implement auto-scaling strategies to handle dynamic workloads and ensure 99.99% uptime.
- Collaborate closely with data scientists to deploy containerized models and MLOps workflows.
- Ensure data sovereignty, security compliance, and resilience in multi-cloud environments.
- Drive the adoption of cutting-edge infrastructure technologies for the 2026 roadmap.
- Conduct performance tuning and cost optimization across the infrastructure stack.
Qualifications
- 5+ years of experience in Systems Engineering, DevOps, or SRE with a focus on AI/ML workloads.
- Proficiency in Python, Kubernetes, Docker, and Terraform.
- Deep understanding of cloud platforms (AWS, GCP, or Azure) and serverless architectures.
- Experience with MLOps tools (MLflow, Kubeflow, Ray) and CI/CD pipelines.
- Strong background in Linux internals, networking, and database performance tuning.
- Bachelor’s degree in Computer Science, Electrical Engineering, or a related technical field.