
Cloud Platform Lead Consultant
Posted 1 day ago

Posted 1 day ago
This is a fully remote position, open to applicants in Texas.
• Design, implement, and manage highly available database platforms such as Apache Cassandra, PostgreSQL, Redis, Valkey, Amazon Redshift, and Google BigQuery across multi-cloud environments.
• Construct, operate, and enhance data streaming infrastructure utilizing Amazon MSK (Kafka), Google Pub/Sub, and Apache Flink to facilitate real-time and batch data pipelines.
• Create and uphold infrastructure-as-code, CI/CD pipelines, and cloud automation through Python and industry-standard tools to ensure repeatable, secure deployments.
• Establish comprehensive monitoring, alerting, and observability for data platform services to proactively identify and resolve issues before they affect customers.
• Collaborate with application development teams to diagnose, fine-tune, and enhance application performance, query patterns, and data access layers supported by team-managed platforms.
• Manage and optimize analytics and query engines, including Starburst Galaxy and AWS Athena, to provide efficient, cost-effective access to large-scale datasets.
• Lead incident response efforts, root cause analysis, and post-incident reviews for production database and streaming systems; drive remediation and preventive enhancements.
• Engage in an on-call rotation to deliver 24x7 support for mission-critical data infrastructure.
• Assess and integrate emerging technologies—including AI agents and MCP servers—to automate operational tasks, enhance developer experience, and expedite DevOps workflows.
• Contribute to capacity planning, disaster recovery, security hardening, and cost optimization initiatives across the data platform landscape.
• 3-5 or more years of overall software engineering or infrastructure experience, with a minimum of 2-4 years in site reliability engineering, DevOps, or platform engineering managing production systems at scale.
• Proven expertise in designing, deploying, and managing cloud infrastructure on AWS and/or Google Cloud Platform, encompassing networking, identity, and security fundamentals.
• Significant hands-on experience with relational and NoSQL databases; production experience with PostgreSQL and at least one distributed database such as Apache Cassandra.
• Experience in operating data streaming platforms; hands-on knowledge of Apache Kafka (including Amazon MSK) and a solid understanding of streaming fundamentals (partitions, consumer groups, delivery semantics, backpressure).
• Advanced skills in production-grade Python and Shell scripting development, including writing, reviewing, and debugging application code, building custom automation tools, and developing operational solutions that extend beyond basic scripting.
• Strong experience with infrastructure-as-code (e.g., Terraform, Terraform Enterprise (TFE), Env0), Jenkins, Ansible, Git CI/CD pipelines, and container orchestration (e.g., Kubernetes) in production environments.
• Background in implementing and automating monitoring, logging, and alerting solutions for distributed systems (e.g., Prometheus, Grafana, CloudWatch, Datadog, or equivalent), including building automated runbooks and self-healing remediation workflows.
• Demonstrated ability to independently lead root cause analysis for complex production incidents that encompass infrastructure, databases, streaming pipelines, and application code layers.
• A comprehensive technology setup, including a laptop, monitors, headset, keyboard, and mouse.
• Monthly connectivity reimbursement to help offset internet costs.
Resolve To Save Lives
Stefanini Brasil
Get handpicked remote jobs straight to your inbox weekly.