Job Description
Are you ready to architect the future of Artificial Intelligence?
Apex Neural Systems is at the forefront of building the computational backbone for the next generation of sentient software. As we prepare to launch our flagship 2026 AI suite, we are seeking a visionary Senior AI Infrastructure Engineer to optimize our high-performance computing (HPC) environments and streamline model deployment pipelines.
In this role, you will bridge the gap between cutting-edge machine learning research and scalable, fault-tolerant production systems. You will work directly with our research leads to ensure our infrastructure can handle the exponential growth of next-gen neural networks.
Why join us? We offer competitive compensation, equity packages, and the chance to define the standard for AI infrastructure in the 2026 landscape.
Responsibilities
- Architect Scalable Systems: Design and implement robust infrastructure for training and deploying large-scale Generative AI models with a focus on speed and efficiency.
- Optimize Compute Resources: Maximize GPU cluster utilization and reduce inference latency through advanced profiling and optimization techniques.
- Cloud & On-Prem Hybrid Strategy: Manage hybrid cloud environments (AWS/Azure/GCP) and on-premise HPC clusters to ensure data sovereignty and cost optimization.
- MLOps Implementation: Build and maintain CI/CD pipelines for machine learning models, automating testing, validation, and deployment processes.
- Reliability Engineering: Develop strategies for system resilience, fault tolerance, and disaster recovery to minimize downtime.
- Team Leadership: Mentor junior engineers and collaborate with data scientists to translate research requirements into engineering solutions.
Qualifications
- Education: Bachelor’s or Master’s degree in Computer Science, Electrical Engineering, or a related technical field.
- Experience: 5+ years of experience in infrastructure engineering, DevOps, or systems architecture within a high-scale tech environment.
- Technical Skills: Deep proficiency in Python, Docker, Kubernetes, and Bash scripting.
- Cloud Expertise: Strong understanding of cloud provider services (AWS, Azure, or GCP) with a focus on AI/ML services (SageMaker, Vertex AI, etc.).
- Performance Tuning: Proven track record of optimizing database performance and system throughput for data-intensive applications.
- Tools: Familiarity with MLflow, Kubeflow, and Prometheus/Grafana for monitoring.