
Software Development Manager, Platform Engineering
Posted 20 hours ago

Posted 20 hours ago
This is a fully remote position, open to applicants in New Jersey.
• Lead, mentor, and facilitate the development of a high-performing Platform/DevOps engineering team.
• Promote an inclusive, collaborative, and psychologically safe team environment.
• Take ownership of delivery for one or more business applications, including executing roadmaps, planning releases, managing risks, and ensuring operational excellence.
• Establish and enhance Service Level Objectives (SLOs), Service Level Indicators (SLIs), on-call practices, incident response workflows, and postmortem culture.
• Collaborate cross-functionally with Product, Design, Security, and Engineering teams to align priorities and effectively communicate progress.
• Manage and advance Kubernetes/EKS environments across multiple Availability Zones (AZs) and regions.
• Champion GitOps and CI/CD best practices utilizing tools such as Argo CD, Flux, GitHub Actions, Jenkins, CodeBuild, and CodeDeploy.
• Create paved roads, templates, and reusable developer workflows to enhance engineering productivity and minimize complexity.
• Facilitate infrastructure automation and infrastructure-as-code methodologies using Terraform and AWS CDK.
• Enhance observability and operational visibility through logs, metrics, tracing, and synthetic monitoring.
• Collaborate with Security and Compliance teams to implement secure-by-default infrastructure practices.
• Assist in optimizing cloud cost efficiency, scalability, reliability, and performance within AWS and Kubernetes ecosystems.
• Foster a culture of continuous improvement, resilience engineering, and operational ownership.
• A minimum of 3 years of experience in people leadership, managing and developing engineering teams.
• At least 7 years of engineering experience in cloud infrastructure, DevOps, platform engineering, or related fields.
• Proven experience in delivering and operating business-critical applications within AWS environments.
• Extensive knowledge of AWS services, including EKS, EC2, VPC networking, S3, RDS/DynamoDB, Lambda, API Gateway, and CloudWatch.
• Expertise in Kubernetes operations, Helm, autoscaling, and infrastructure security practices.
• Experience in constructing CI/CD pipelines and GitOps workflows.
• Strong understanding of Infrastructure as Code principles utilizing Terraform and/or AWS CDK.
• Proficient in programming languages such as Python, Java, TypeScript, or Node.js.
• Familiarity with Site Reliability Engineering (SRE) practices, including incident management, SLIs/SLOs, and operational excellence.
• Excellent communication skills and ability to partner with stakeholders.
• Experience in building or managing Internal Developer Platforms and self-service engineering workflows.
• Knowledge of service meshes, policy engines, secrets management, and platform governance.
• Experience in supporting multi-region architectures and resilience engineering initiatives.
• Understanding of observability platforms, distributed tracing, and synthetic monitoring strategies.
• Experience with event-driven architectures and messaging systems such as Kafka, Kinesis, SNS, or SQS.
• Familiarity with SOX, SOC2, compliance frameworks, and FinOps practices.
• Health benefits.
• Financial benefits.
• Time away.
• Everyday wellness benefits.
• Annual cash bonuses.
• Stock grants.
• Comprehensive benefits package.
• In-person onboarding and/or in-person ID verification may be provided.
GoMining
Calliere Group
CData Software
URUS Group
Get handpicked remote jobs straight to your inbox weekly.