
MLOps Engineer
Posted 1 day ago

Posted 1 day ago
This is a fully remote position, open to applicants in India.
• Take charge of the infrastructure that underpins production AI systems.
• Transform AI services from single-host setups to horizontally scalable, orchestrated infrastructures.
• Create infrastructure tailored for LLM workloads, accommodating long-running requests, streaming responses, fluctuating concurrency, costly downstream calls, and upstream rate limitations.
• Manage capacity, autoscaling, and unit economics effectively.
• Establish promotion pathways from development to pre-production and production, ensuring consistent and reproducible environments.
• Implement version control to make infrastructure, configuration, and application logic subject to review.
• Develop CI/CD processes that facilitate rapid deployment, rollback, staged rollouts, and maintain an auditable change history.
• Create self-service tools enabling engineers and technically inclined colleagues to define, modify, and test AI workflow logic.
• Design validation, versioning, review, staged promotion, and recovery safeguards.
• Monitor the AI stack for latency, throughput, failure modes, cost, and output quality.
• Establish effective alerting and incident management practices to enhance the system.
• Manage secrets, access control, and environment isolation within a regulated sector.
• Proven experience managing production infrastructure: deployed and operated containerized services in live environments, including rollouts, autoscaling, failure isolation, resource limits, and rollback procedures.
• Proficiency in infrastructure as code; capable of creating declarative and reproducible environments.
• Familiarity with CI/CD practices, encompassing automated testing, environment promotion, safe rollouts, and quick recovery.
• Strong skills in Python; able to read, modify application code, profile, and troubleshoot it.
• Solid operational judgment spanning application, network, and infrastructure domains.
• Comprehensive ownership of design, implementation, deployment, monitoring, and follow-up processes.
• Capacity for fast, incremental iteration and risk mitigation.
• Experience in managing LLM or ML workloads in production settings.
• Cloud deployment experience, preferably with AWS.
• Background in fintech, lending, insurance, or another regulated industry.
• Experience in developing internal developer platforms or self-service tools.
• Comfort with full-stack development for lightweight UIs or internal tools.
• Experience in designing multi-environment promotion pipelines.
• Familiarity with agent orchestration frameworks.
• Knowledge of observability practices for non-deterministic systems.
• Skills in inference optimization, including model serving, batching, caching, and cost reduction strategies.
• Exposure to workflow automation tools and low-code development platforms.
• Competitive compensation package.
• Unlimited paid time off (PTO).
• Remote-first work environment with flexible hours.
• Annual professional development budget of $2,000.
• Stipend for home office setup.
Sourcegraph
Quora
NBCUniversal
Spotify
Get handpicked remote jobs straight to your inbox weekly.