
Cloud Platform Lead Consultant
Posted 1 day ago

Posted 1 day ago
This is a fully remote position, open to applicants in United States.
• Build, manage, and enhance data streaming infrastructure utilizing Amazon MSK (Kafka) or Google Pub/Sub.
• Design, implement, and oversee highly available database and caching platforms across multi-cloud settings.
• Create and sustain infrastructure-as-code, CI/CD pipelines, and cloud automation using Terraform and Python.
• Establish monitoring, alerting, and observability for data platform services.
• Collaborate with application development teams to diagnose, optimize, and enhance application performance and data access layers.
• Administer and refine Starburst Galaxy and AWS Athena.
• Engage in incident response, root cause analysis, and post-incident reviews.
• Participate in an on-call rotation to support critical data infrastructure.
• Review application source code and work together on targeted fixes and optimization strategies.
• Assess emerging tools and AI-driven automation methods.
• Contribute to capacity planning, disaster recovery, security enhancements, and cost efficiency.
• 3–5 years of experience in software engineering or infrastructure, with a minimum of 2 years in SRE, DevOps, or platform engineering managing production systems at scale.
• Practical experience in designing, deploying, and managing cloud infrastructure on AWS and/or Google Cloud Platform, including networking, identity, and security fundamentals.
• Hands-on experience with data streaming platforms; familiarity with Apache Kafka, including Amazon MSK or Confluent Kafka, and knowledge of partitions, consumer groups, delivery semantics, and backpressure.
• Production experience with both relational and NoSQL databases; PostgreSQL is a must.
• Proficient in Python and Shell scripting for automation and operational solutions.
• Extensive experience with infrastructure-as-code tools like Terraform, CI/CD tools such as Jenkins and Git, Ansible, and container orchestration platforms like Kubernetes in production settings.
• Experience in implementing and automating monitoring, logging, and alerting for distributed systems.
• Demonstrated ability to contribute to root cause analysis for production incidents.
• Strong problem-solving, communication, and documentation skills with accountability in on-call and incident management scenarios.
• A solid understanding of distributed systems principles, including high availability, fault tolerance, consistency models, and disaster recovery.
• Background investigation is required.
• A dedicated, private workspace and reliable internet connection with minimum speeds of 50 MB download and 5 MB upload while working from home.
• Comprehensive technology setup that includes a laptop, monitors, headset, keyboard, and mouse.
• Monthly connectivity reimbursement to help offset internet costs for employees eligible to work from home.
Public Consulting Group
Allstate
Samsara
Guidehouse
Get handpicked remote jobs straight to your inbox weekly.