
Senior Software Engineer β Infra Agent Systems
Posted Aug 26

Posted Aug 26
This is a fully remote position, open to applicants in India.
β’ Design and develop production AI agent systems that diagnose, investigate, and resolve infrastructure issues across one of the largest GPU fleets globally.
β’ Create distributed services, orchestration frameworks, knowledge graphs, and retrieval systems that empower infrastructure agents.
β’ Build fleet intelligence systems that integrate telemetry, infrastructure state, operational knowledge, and historical incidents.
β’ Connect with observability, incident management, ticketing, fleet inventory, source control, chat, and internal infrastructure systems via APIs.
β’ Take ownership of services from start to finish, encompassing architecture, implementation, testing, deployment, observability, and production operations.
β’ Enhance agent performance through evaluations, retrieval enhancements, improved tools, and production feedback loops.
β’ Transform agentsβ production insights into robust, reviewed software and automation.
β’ Deliver, operate, and provide support for software in a production environment.
β’ Over 5 years of experience in building production backend systems, distributed systems, or infrastructure platforms.
β’ Strong skills in systems design and experience in managing significant systems from design through to production.
β’ Expertise in at least one of the following areas: AI agent systems, orchestration, tool utilization, evaluation, or grounding; knowledge graphs or graph data modeling; search, retrieval, ranking, RAG, or semantic search systems.
β’ Proficient in backend engineering, particularly in API design, service boundaries, data modeling, and integrations within complex systems.
β’ Familiarity with Kubernetes, GitOps tools like ArgoCD, infrastructure-as-code, and cloud platforms.
β’ Comfortable programming in Go, TypeScript, Python, or Rust.
β’ Knowledge of GPU infrastructure, datacenters, bare-metal systems, hardware failure modes, BMC/IPMI, or cluster schedulers is advantageous.
β’ Experience with graph databases is a plus.
β’ Familiarity with event-driven systems and messaging platforms such as NATS or Kafka is a plus.
β’ Experience with observability platforms like Prometheus and Grafana is a plus.
β’ Background in building evaluation frameworks or enhancing the quality and reliability of LLM-powered systems is a plus.
β’ Competitive salary and performance-based bonuses.
β’ Comprehensive health, dental, and vision insurance.
β’ Flexible working hours and remote work options.
β’ Opportunities for professional development and training.
β’ Collaborative and innovative work environment.
RTX
NVIDIA
Infios
Gramian Consulting
Get handpicked remote jobs straight to your inbox weekly.