Job Description
We are seeking a visionary Senior AI Infrastructure Architect to join our elite team and lead the technical vision for Project 2026. As we push the boundaries of artificial general intelligence, you will be responsible for designing the scalable, fault-tolerant, and high-performance computing infrastructure that powers our next-generation neural networks.
In this role, you will bridge the gap between cutting-edge AI research and production-grade engineering. You will optimize data pipelines, manage massive-scale GPU clusters, and ensure our systems are secure, efficient, and ready for the demands of a future AI-driven economy.
Why Join Us?
- Work on the bleeding edge of AI technology.
- Competitive equity package and performance bonuses.
- Top-tier benefits and remote-first flexibility.
Responsibilities
- System Design: Architect and implement a robust distributed computing infrastructure optimized for large-scale deep learning model training and inference.
- Performance Optimization: Identify bottlenecks in data pipelines and hardware utilization, implementing solutions to reduce training time and improve model accuracy.
- Cluster Management: Oversee the deployment and management of high-performance GPU clusters using Kubernetes and container orchestration technologies.
- Reliability Engineering: Build fault-tolerant systems that ensure 99.99% uptime for critical AI workloads.
- Cross-Functional Leadership: Collaborate with research scientists and software engineers to translate theoretical models into production-ready software.
- Security & Compliance: Implement rigorous security protocols to protect sensitive data and intellectual property.
Qualifications
- Education: Bachelor’s or Master’s degree in Computer Science, Electrical Engineering, or a related field (PhD preferred).
- Experience: 7+ years of experience in systems architecture, with at least 3 years specifically focused on AI/ML infrastructure.
- Programming: Proficiency in Python, C++, and CUDA for high-performance computing.
- Tools: Deep expertise in Kubernetes, Docker, AWS/GCP, and Apache Spark or similar data processing frameworks.
- Soft Skills: Exceptional problem-solving abilities and the ability to communicate complex technical concepts to non-technical stakeholders.
- Passion: A deep passion for the future of AI and a desire to solve humanity's hardest problems.