
Senior Platform Engineer
Posted Jul 17

Posted Jul 17
This is a fully remote position, open to applicants in India.
• Create and enhance cloud-native platform architecture that accommodates multi-tenant SaaS applications.
• Establish infrastructure standards, deployment methodologies, and best practices for the platform.
• Lead architecture assessments and analyze technical trade-offs involving reliability, performance, security, and cost.
• Design systems that are highly available and fault-tolerant across various environments.
• Architect and oversee extensive AWS environments.
• Develop networking architectures, including VPCs, subnets, security groups, routing, load balancing, and connectivity strategies.
• Construct secure deployment architectures that meet security and compliance standards.
• Execute disaster recovery, backup, and business continuity plans.
• Design and manage production Kubernetes environments.
• Formulate scalable container orchestration strategies.
• Enhance cluster performance, networking, autoscaling, and workload scheduling.
• Elevate developer experience through platform automation and self-service tools.
• Provide support for AI and ML workloads operating on AWS.
• Create infrastructure for model training and inference tasks.
• Oversee GPU provisioning, utilization, scaling, and cost management.
• Collaborate with AI teams to enhance deployment and operational efficiency.
• Define and track SLIs, SLOs, and operational metrics.
• Implement practices for monitoring, observability, logging, alerting, and incident management.
• Lead initiatives for performance optimization and capacity planning.
• Conduct root cause analysis and spearhead reliability enhancement efforts.
• Develop Infrastructure-as-Code solutions utilizing Terraform.
• Design and refine CI/CD pipelines.
• Automate provisioning, deployments, scaling, and operational procedures.
• Continuously assess cloud expenditures.
• Formulate capacity planning models.
• Balance performance, reliability, and infrastructure costs.
• 6-8 years of experience in DevOps, Platform Engineering, SRE, or Cloud Infrastructure roles.
• Demonstrated experience in designing and managing production-scale SaaS platforms.
• Strong knowledge of AWS architecture, networking, security, and deployment strategies.
• Extensive hands-on experience with Kubernetes, container orchestration, cluster operations, autoscaling, and workload management.
• Proven expertise in designing highly available, fault-tolerant, and scalable distributed systems.
• Solid understanding of system design, architecture trade-offs, and platform scalability.
• Direct experience with Infrastructure as Code (preferably Terraform).
• Familiarity with building and maintaining CI/CD pipelines and deployment automation frameworks.
• Strong fundamentals in Linux, networking, and systems engineering.
• Experience in implementing observability, monitoring, logging, and incident management practices.
• Knowledge of disaster recovery planning, backup strategies, and business continuity design.
• Experience with cloud cost optimization, capacity planning, and resource utilization management.
• Hands-on experience supporting AI/ML workloads in production settings.
• Experience in designing, provisioning, and managing GPU-based infrastructure for model training and/or inference tasks.
• Familiarity with managing and optimizing AWS Bedrock, OpenSearch, and DocumentDB or equivalent platforms.
• Strong scripting and automation capabilities using Python, Bash, or similar languages.
• Competitive salary and performance-based bonuses.
• Comprehensive health, dental, and vision insurance.
• Flexible work hours and remote work options.
• Opportunities for professional development and continuous learning.
• Supportive and dynamic team environment.
Arctiq
Cisco
Prove
Hello Heart
Get handpicked remote jobs straight to your inbox weekly.