Senior Software Engineer – Infra Agent Systems

Posted Aug 26

This is a fully remote position, open to applicants in India.

πŸ“‹ Description

β€’ Design and develop production AI agent systems that diagnose, investigate, and resolve infrastructure issues across one of the largest GPU fleets globally.

β€’ Create distributed services, orchestration frameworks, knowledge graphs, and retrieval systems that empower infrastructure agents.

β€’ Build fleet intelligence systems that integrate telemetry, infrastructure state, operational knowledge, and historical incidents.

β€’ Connect with observability, incident management, ticketing, fleet inventory, source control, chat, and internal infrastructure systems via APIs.

β€’ Take ownership of services from start to finish, encompassing architecture, implementation, testing, deployment, observability, and production operations.

β€’ Enhance agent performance through evaluations, retrieval enhancements, improved tools, and production feedback loops.

β€’ Transform agents’ production insights into robust, reviewed software and automation.

β€’ Deliver, operate, and provide support for software in a production environment.


⛳️ Requirements

β€’ Over 5 years of experience in building production backend systems, distributed systems, or infrastructure platforms.

β€’ Strong skills in systems design and experience in managing significant systems from design through to production.

β€’ Expertise in at least one of the following areas: AI agent systems, orchestration, tool utilization, evaluation, or grounding; knowledge graphs or graph data modeling; search, retrieval, ranking, RAG, or semantic search systems.

β€’ Proficient in backend engineering, particularly in API design, service boundaries, data modeling, and integrations within complex systems.

β€’ Familiarity with Kubernetes, GitOps tools like ArgoCD, infrastructure-as-code, and cloud platforms.

β€’ Comfortable programming in Go, TypeScript, Python, or Rust.

β€’ Knowledge of GPU infrastructure, datacenters, bare-metal systems, hardware failure modes, BMC/IPMI, or cluster schedulers is advantageous.

β€’ Experience with graph databases is a plus.

β€’ Familiarity with event-driven systems and messaging platforms such as NATS or Kafka is a plus.

β€’ Experience with observability platforms like Prometheus and Grafana is a plus.

β€’ Background in building evaluation frameworks or enhancing the quality and reliability of LLM-powered systems is a plus.


🏝️ Benefits

β€’ Competitive salary and performance-based bonuses.

β€’ Comprehensive health, dental, and vision insurance.

β€’ Flexible working hours and remote work options.

β€’ Opportunities for professional development and training.

β€’ Collaborative and innovative work environment.

People also viewed

RTX15 hours ago

Principal Software Engineer – Air Traffic Solutions

US flagMaine OnlyFull-timeFull-stack Engineer$107.5k – $204.5k/year
ApplyView job
NVIDIA16 hours ago

Senior System Software Engineer, Software-Defined Networking

US flagCalifornia, +4 more statesFull-timeFull-stack Engineer$224k – $356.5k/year
ApplyView job
Infios17 hours ago

Senior Software Engineer

MX flagMexico, +1 more countryFull-timeFull-stack Engineer
ApplyView job
Gramian Consulting1 day ago

Software Engineer – Licensing, AI Training

EG flagEgypt, +3 more countriesFull-timeFull-stack Engineer$100/year
ApplyView job
Gramian Consulting1 day ago

Software Engineer – Licensing, AI Training

IN flagIndia, +5 more countriesFreelanceFull-stack Engineer$100/year
ApplyView job
Gramian Consulting1 day ago

Software Engineer – License your Git Repositories for AI Training

CA flagCanada, +5 more countriesFull-timeFull-stack Engineer$100/year
ApplyView job

Never miss a great job!

Get handpicked remote jobs straight to your inbox weekly.

Trusted by 7,400+ designers