
Senior Platform Engineer
Posted Jul 30

Posted Jul 30
This is a fully remote position, open to applicants in India.
• Create and enhance cloud-native platform architecture for multi-tenant SaaS applications
• Establish infrastructure standards, deployment patterns, and best practices for the platform
• Lead architecture assessments and analyze technical trade-offs related to reliability, performance, security, and cost
• Design systems that are highly available and fault-tolerant across various environments
• Architect and oversee large-scale AWS infrastructures
• Develop networking architectures that include VPCs, subnets, security groups, routing, load balancing, and connectivity patterns
• Construct secure deployment architectures that comply with security and regulatory requirements
• Implement strategies for disaster recovery, backup, and business continuity
• Design and manage production Kubernetes environments
• Develop scalable strategies for container orchestration
• Enhance cluster performance, networking, autoscaling, and workload scheduling
• Improve the developer experience through platform automation and self-service tools
• Support AI and ML workloads hosted on AWS
• Design infrastructure suitable for model training and inference tasks
• Oversee GPU provisioning, utilization, scaling, and cost management
• Collaborate with AI teams to enhance deployment and operational efficiency
• Define and monitor SLIs, SLOs, and operational metrics
• Implement practices for monitoring, observability, logging, alerting, and incident management
• Lead initiatives for performance optimization and capacity planning
• Conduct root cause analyses and drive reliability improvement efforts
• Build Infrastructure-as-Code solutions using Terraform
• Design and refine CI/CD pipelines
• Automate workflows for provisioning, deployments, scaling, and operations
• Continuously assess cloud spending
• Develop models for capacity planning
• Balance performance, reliability, and infrastructure expenses
• 6-8 years of experience in DevOps, Platform Engineering, SRE, or Cloud Infrastructure roles
• Demonstrated experience in designing and operating production-scale SaaS platforms
• Strong proficiency in AWS architecture, networking, security, and deployment methodologies
• Extensive hands-on experience with Kubernetes, container orchestration, cluster operations, autoscaling, and workload management
• Experience in designing systems that are highly available, fault-tolerant, and scalable
• Comprehensive understanding of system design, architectural trade-offs, and platform scalability
• Practical experience with Infrastructure as Code (Terraform preferred)
• Background in building and maintaining CI/CD pipelines and deployment automation frameworks
• Strong fundamentals in Linux, networking, and systems engineering
• Familiarity with implementing observability, monitoring, logging, and incident management protocols
• Experience in disaster recovery planning, backup strategies, and business continuity design
• Knowledge in cloud cost optimization, capacity planning, and resource utilization management
• Practical experience supporting AI/ML workloads in production settings
• Expertise in designing, provisioning, and managing GPU-based infrastructure for model training and/or inference tasks
• Experience with managing and optimizing AWS Bedrock, OpenSearch, and DocumentDB or similar platforms
• Proficient in scripting and automation using Python, Bash, or similar languages
• Competitive salary and performance-based bonuses
• Comprehensive health, dental, and vision insurance
• Flexible work hours and remote work options
• Opportunities for professional development and growth
• Collaborative and innovative work environment
Quantiphi
Encompass Corporation
Shippit
Group O
Get handpicked remote jobs straight to your inbox weekly.