
Senior Platform Engineer
Posted 19 hours ago

Posted 19 hours ago
This is a fully remote position, open to applicants in Romania.
• Design, construct, and manage AWS infrastructure as code utilizing Terraform and Terragrunt across a multi-account setup and around 70 service repositories.
• Contribute to the AWS multi-account initiative, which includes Organizations, IAM Identity Center, Service and Resource Control Policies, account bootstrap, and baselines.
• Develop and maintain shared CI/CD and release tools using GitHub Actions, AWS CodeBuild, CodeDeploy, semantic-release, and policy-as-code checks.
• Manage containerized and serverless production workloads on ECS Fargate and Lambda.
• Oversee Datadog monitoring, logging, and alerting; design actionable alerts and manage observability costs.
• Participate in the weekly support rota and Datadog On-Call; respond to incidents and conduct root cause analyses.
• Utilize security tools and drive vulnerability findings through remediation.
• Manage AWS WAF, CloudFront, firewalls, client IP whitelisting, and certificate lifecycles.
• Support ISO 27001, DORA, and SWIFT compliance, including annual penetration testing and disaster recovery exercises.
• Operate and upgrade Aurora MySQL, RDS Proxy, DocumentDB, Redis, and Amazon MQ.
• Deliver client-facing infrastructure onboarding, which includes SFTP, AWS Transfer Family, key management, tenant, and sandbox provisioning.
• Manage secrets, identity, and access across AWS and related systems.
• Contribute to FinOps activities, including cost analysis, budgeting, guardrails, rightsizing, and backup lifecycle management.
• Automate manual runbook tasks through self-service tools.
• Modernize and safely decommission legacy systems.
• Document work in runbooks, ADRs, and threat models.
• Operate within Agile methodologies, including two-week sprints and backlog refinement.
• Share knowledge and provide support to less experienced engineers in the distributed team.
• A minimum of 5 years of hands-on experience in Platform Engineering, DevOps, SRE, or Infrastructure roles, or an equivalent.
• At least 3 years of direct experience designing and managing production workloads on AWS.
• A minimum of 2 years managing production environments on AWS ECS, Fargate, and serverless Lambda.
• Proven experience writing and maintaining production Terraform at scale.
• Experience with or a willingness to work with Terragrunt or an equivalent multi-account orchestration tool.
• Familiarity with AWS multi-account architecture, including AWS Organizations, IAM Identity Center, Service Control Policies, and account baselining.
• Experience in production on-call duties and incident response in a customer-facing setting.
• Hands-on experience with vulnerability management and remediation.
• Experience working within a recognized security or regulatory framework such as ISO 27001, SOC 2, or DORA.
• Proficiency in reading and debugging Node.js services.
• Ability to write production-quality code in Python and Bash.
• Experience in financial institutions or regulated SaaS environments is highly desirable.
• In-depth knowledge of AWS services, including ECS/Fargate, Lambda, EC2, RDS/Aurora MySQL, DocumentDB, Redis, Amazon MQ, S3, EFS, ELB/ALB, Route 53, API Gateway, CloudFront, WAF, ACM, Secrets Manager, SSM Parameter Store, KMS, AWS Backup, ECR, CodeBuild, AWS Transfer Family, Organizations, and IAM Identity Center.
• Expertise in Terraform, including module design, remote state management, and workspace strategy.
• Proficient in using Ansible.
• Experience implementing CI/CD pipelines with GitHub Actions and AWS CodeBuild/CodeDeploy.
• Familiarity with Datadog for observability, metrics, logs, APM, database monitoring, and on-call scheduling.
• Experience with security tools such as Wazuh, Prowler, Trivy, Checkov, or similar.
• Strong understanding of networking and system architecture, including VPCs, peering, endpoints, DNS, TLS, certificates, zero-trust access, and firewalls.
• Linux administration experience on Ubuntu; experience with CentOS is a plus.
• Knowledge of Nginx and Apache.
• Proficiency in Python, Bash, HCL, YAML, and SQL.
• Understanding of microservice and event-driven architecture, including SQS, SNS, EventBridge, RESTful APIs, and SPAs.
• Familiarity with FinOps practices.
• Experience using Confluence and Jira.
• Knowledge and application of Agile methodologies.
• Ability to work during production incidents and scheduled weekend change windows.
• Excellent written and verbal communication skills across time zones and functions.
• Remote work opportunity in Romania.
• Supportive distributed team environment.
• Opportunities to collaborate with senior leadership and cross-functional teams.
• Exposure to advanced AWS, SaaS, FinTech, security, compliance, and FinOps practices.
• Knowledge sharing and support for professional development.
• Two-week Agile sprints and collaborative engineering practices.
Pax8
CVS Health
Bright Machines
Vitalize Unlimited
Get handpicked remote jobs straight to your inbox weekly.