
Senior AI Software Engineer, Python
Posted Jun 24

Posted Jun 24
This is a fully remote position, open to applicants in United States.
• Take charge of the reliability of the event-driven messaging layer, which includes managing backpressure, ensuring idempotency, handling dead-letter scenarios, and implementing retry strategies.
• Design and manage the infrastructure necessary for executing LLM orchestration workloads at scale.
• Oversee the operational data layer for the CI runtime, focusing on state management, session persistence, and real-time data access patterns.
• Ensure observability for the CI platform through structured logging, distributed tracing (OpenTelemetry), and error tracking (Sentry).
• Maintain and strengthen the interfaces between CI and downstream platforms, which involves contract testing, versioning, and failure management.
• Conduct code reviews and provide mentorship to team members on best practices in Python engineering and production readiness.
• Manage production support for CI infrastructure, which includes on-call responsibilities and incident response.
• 5+ years of experience or relevant expertise.
• Proven track record in developing production-grade Python services at scale.
• Experience operating real-time or high-throughput infrastructure that supports AI/ML workloads.
• Skilled in designing and managing distributed, event-driven systems in production settings.
• Profound understanding of Python internals, including async/await lifecycle, event loop mechanics, GIL implications for concurrency strategies, and memory profiling.
• Hands-on experience with Pydantic, type systems, and structured data modeling in high-throughput services.
• Strong views on code organization, error handling practices, and testability within long-lived Python codebases.
• Practical experience operating LLM inference infrastructure at scale.
• Extensive background in NoSQL data modeling, including partition strategies, consistency trade-offs, query cost optimization, and avoiding hot partitions.
• Familiarity with event-driven architecture in production, including backpressure management, idempotency, dead-letter handling, and retry strategies.
• Proficient with observability tools, such as distributed tracing (OpenTelemetry), structured logging, and error tracking (Sentry).
• Experience with services on the Azure cloud platform.
• 100% employer-covered medical, dental, and vision insurance.
• Flexible paid time off, allowing you to rest, relax, and recharge away from work.
• Paid parental leave.
• Compensation for cell phone and service costs.
• Remote office allowance.
• Opportunities for professional development and training courses.
• Regular check-ins with managers to enhance performance and career growth through Lattice.
One Identity
9th Way Insignia
Engineering System and Technologies SARL
Get handpicked remote jobs straight to your inbox weekly.