
Software Engineer, Infrastructure, Reliability
Posted 1 day ago

Posted 1 day ago
This is a fully remote position, open to applicants in United States.
• Take ownership of and enhance the infrastructure that supports CrewAI's platform, which encompasses AWS, ECS/ECR, Docker, Kubernetes/Helm, networking, secrets, databases, Redis, and related services.
• Develop and sustain CI/CD pipelines for building, testing, image publishing, migrations, environment promotion, rollbacks, and ensuring deployment safety.
• Enhance reliability across both cloud and enterprise deployments through health checks, alerting, incident response, capacity planning, recovery paths, and operational runbooks.
• Lead the front-line on-call rotation and uphold its SLAs.
• Collaborate with runtime engineers on Celery/FastAPI/Redis workloads and product engineers on Rails/Solid Queue/Postgres production performance.
• Oversee production observability and telemetry infrastructure, which includes logs, metrics, traces, dashboards, Sentry/OpenTelemetry integration, actionable alerts, and telemetry export to clients' monitoring systems.
• Strengthen security and compliance across IAM, workload identity, secrets management, vulnerability scanning, dependency/image hygiene, and least-privilege access.
• Create tools and automation for field engineers and customers to conduct self-hosted installations, including Helm charts, environment configuration, release artifacts, pre-flight checks, and install runbooks.
• Mitigate operational toil by automating repetitive workflows and streamlining deployments.
• Proven infrastructure/platform engineering experience in production SaaS environments.
• Extensive practical knowledge of AWS, Docker, CI/CD, GitHub Actions, and containerized services.
• Familiarity with ECS and/or Kubernetes; experience with Helm is highly desirable.
• Proficient in managing PostgreSQL, Redis, background job systems, queues, and web services in a production setting.
• Strong debugging skills across application, infrastructure, network, deployment, and dependency layers.
• Security-focused mindset regarding IAM, secrets, workload identity, vulnerability management, and production access.
• Capability to write reliable automation scripts in Python, Ruby, Go, Bash, or similar languages.
• Composed and meticulous approach to incidents, rollbacks, migrations, and production change management.
• Bonus: Experience with AI/agent platforms, workflow runtimes, or high-volume asynchronous execution systems.
• Bonus: Background in supporting enterprise/self-hosted deployments.
• Bonus: Familiarity with Terraform or other Infrastructure as Code (IaC) tools.
• Bonus: Experience in Site Reliability Engineering (SRE) including SLOs, incident review, capacity planning, and load testing.
• Bonus: Knowledge of Rails, FastAPI, Celery, OpenTelemetry, or multi-service observability.
• Comprehensive health insurance coverage.
• Flexible work hours and the possibility of remote work.
• Opportunities for professional development and career advancement.
• A collaborative and innovative work environment.
VALR
First Circle
Sezzle
Sezzle
Get handpicked remote jobs straight to your inbox weekly.