Remotery

Senior DevOps Engineer – AWS, AI Infrastructure

Posted Jun 13

This is a fully remote position, open to applicants in Argentina.

📋 Description

• Provision and set up a dedicated VPC along with a segmented cloud environment on AWS.

• Establish the foundational CI/CD pipeline and oversee its maintenance and evolution throughout all delivery stages.

• Configure and manage the vector store infrastructure (OpenSearch/Pinecone on AWS).

• Implement and manage the observability stack, including CloudWatch, X-Ray, alerting thresholds, and monitoring specific to LLM.

• Execute infrastructure-as-code for all environments (development, staging, production) utilizing Terraform or CDK.

• Oversee secrets management, KMS encryption key configuration, and tenant-scoped access controls.

• Set up connectivity with LLM providers (OpenAI / Anthropic / Amazon Bedrock enterprise tier, zero-data-retention).

• Develop and execute an environment promotion strategy in line with the 2-week sprint cadence.

• Assist with the infrastructure requirements for the incremental ingestion pipeline and manage nightly scheduling.


⛳️ Requirements

• Over 6 years of experience in DevOps or cloud infrastructure engineering, with a strong emphasis on AWS.

• Proficiency in infrastructure-as-code tools: Terraform, CloudFormation, or AWS CDK.

• Familiarity with CI/CD tools: GitHub Actions, AWS CodePipeline, or similar.

• Knowledge of core AWS services: VPC, ECS, Lambda, S3, DynamoDB, API Gateway, Cognito, CloudWatch, X-Ray.

• Proven experience in designing and managing multi-tenant cloud environments with tenant-level data isolation.

• AI experience is required and not optional.

• Experience in configuring and managing vector store infrastructure (OpenSearch, Pinecone, Weaviate, or equivalent) in a production setting.

• Understanding of LLM provider APIs (OpenAI, Anthropic, or Amazon Bedrock) in a production or enterprise configuration, including zero-data-retention tier setup.

• Comprehension of AI-specific observability issues: token usage monitoring, latency profiling for LLM calls, and model response logging.


🏝️ Benefits

• An excellent work environment certified by Great Place To Work.

• Opportunities for professional development.

People also viewed

The CodestJul 26

DevOps Engineer

PL flagPoland OnlyFull-timeDevOps & Site Reliability Engineer (SRE)
ApplyView job
IRIUMJul 26

Ingeniero/a Cloud DevOps

ES flagSpain OnlyFull-timeDevOps & Site Reliability Engineer (SRE)€33k – €40k/year
ApplyView job
SólidesJul 26

Senior DevOps Engineer

BR flagBrazil OnlyFull-timeDevOps & Site Reliability Engineer (SRE)
ApplyView job
ResilincJul 25

Junior/Senior Site Reliability Engineer – Night Shift

IN flagIndia OnlyFull-timeDevOps & Site Reliability Engineer (SRE)
ApplyView job
Verity GroupJul 25

Senior SRE / DevOps Engineer

Anywhere in the WorldFull-timeDevOps & Site Reliability Engineer (SRE)
ApplyView job
HOESSLER & HOESSLERJul 25

DevOps Software Engineer – Career Ambitions

DE flagGermany OnlyFull-timeDevOps & Site Reliability Engineer (SRE)€65k – €75k/year
ApplyView job

Never miss a great job!

Get handpicked remote jobs straight to your inbox weekly.

Trusted by 7,400+ designers