
Principal Cloud Engineer
Posted 4 days ago

Posted 4 days ago
This is a fully remote position, open to applicants in United States.
• Design, implement, enhance, and provide operational support for AWS-based cloud environments and hybrid cloud infrastructures.
• Develop and uphold cloud architecture patterns, reference architectures, reusable templates, and engineering standards.
• Lead intricate cloud initiatives spanning infrastructure, application support, data engineering, cybersecurity, networking, and operations.
• Design, deploy, document, and validate environments for development, testing, staging, production, and disaster recovery.
• Configure and manage AWS services to fulfill business, mission, security, availability, and performance needs.
• Oversee Amazon VPC environments, IAM practices, infrastructure as code, provisioning, configuration, scaling, monitoring, remediation, and operational workflows.
• Implement observability measures using Amazon CloudWatch along with associated monitoring, logging, alerting, tracing, and dashboarding tools.
• Conduct incident response, root-cause analysis, problem management, capacity planning, performance tuning, and reliability enhancements.
• Incorporate security-by-design, vulnerability management, encryption, secrets management, audit readiness, and compliance alignment.
• Integrate cloud and on-premises environments, covering connectivity, identity, data movement, backups, storage, and operational support.
• Provide support for Linux and Windows environments, clustered file systems, storage, networking, SAN, backup, and large-scale data platforms.
• Execute technical upgrades, lifecycle management, migrations, and containerization of legacy workloads.
• Create architecture diagrams, engineering standards, implementation plans, runbooks, system design documents, operating procedures, and knowledge articles.
• Perform proofs of concept and technical evaluations for cloud, automation, data, security, and AI technologies.
• Develop and implement system, integration, performance, security, resiliency, and disaster recovery testing.
• Provide technical mentorship, conduct code and design reviews, and offer guidance to cloud engineers, systems administrators, and DevOps practitioners.
• Communicate technical findings, risks, dependencies, trade-offs, and recommendations to management, customers, and technical stakeholders.
• Enhance engineering efficiency, platform reliability, customer experience, security posture, automation, and cloud cost management.
• Utilize AI-assisted engineering workflows and assess approved tools such as Claude Code and Amazon Q Developer.
• Support AI and data workloads utilizing platforms like Amazon Bedrock and Amazon SageMaker.
• Implement responsible AI guardrails and leverage AIOps for event correlation, anomaly detection, predictive capacity planning, and automated remediation.
• A Bachelor's degree in Computer Science, Information Technology, Engineering, or a related field; an equivalent combination of education, technical training, and relevant experience may be accepted.
• 6–10 years of progressive experience in cloud engineering, systems engineering, DevOps, infrastructure engineering, site reliability engineering, or a related technical field.
• Practical experience in designing, deploying, administering, and supporting AWS cloud environments in production settings.
• Strong expertise with AWS services across compute, networking, identity, storage, monitoring, security, and automation domains.
• Experience in designing and managing AWS networking, which includes VPCs, subnets, route tables, security groups, network ACLs, private connectivity, load balancers, DNS, and firewall integration.
• In-depth knowledge of AWS IAM, access control, authentication and authorization, least privilege, encryption, secrets management, and cloud security best practices.
• Hands-on experience with infrastructure as code using AWS CloudFormation, Terraform, or similar tools.
• Experience in implementing CI/CD pipelines and automated deployment practices.
• Familiarity with monitoring, logging, alerting, and operational diagnostics using Amazon CloudWatch and related observability tools.
• Proficiency in Python, Bash, PowerShell, and JavaScript/Node.js.
• Experience administering or supporting Linux and Windows operating systems in enterprise or production environments.
• Strong networking knowledge, including TCP/IP, DNS, TLS, routing, firewall concepts, load balancing, VPNs, network segmentation, and traffic inspection architectures.
• Experience supporting databases and data platforms such as MySQL, PostgreSQL, Oracle, SQL Server, Snowflake, or similar technologies.
• Experience in a 24/7 production environment including operational support, incident response, root-cause analysis, and change management.
• Ability to document complex technical solutions clearly and effectively.
• Capability to convert mission, business, security, and operational requirements into technical designs and implementation plans.
• Experience applying secure software development, DevSecOps, and automation practices in an Agile or iterative delivery environment.
• Familiarity with enterprise-approved generative AI tools and engineering productivity use cases.
• Ability to evaluate AI-generated code, scripts, infrastructure configurations, and technical recommendations for accuracy, security, maintainability, and operational suitability.
• AWS certifications and additional experience as outlined under desired qualifications are preferred but not mandatory.
• Comprehensive benefits and wellness packages.
• 401(k) with company match.
• Competitive compensation.
• Paid time off.
• AI-powered career tool that identifies career steps and learning opportunities.
• Internal mobility team dedicated to career goals.
• Flexible work weeks.
• Award-winning culture of innovation.
• Military-friendly workplace.
• Medical plan options, some of which include Health Savings Accounts.
• Dental plan options.
• Vision plan options.
• Leave options including vacation, sick, personal, holiday, paid parental, military, bereavement, and jury duty.
• Typically, 15 days of paid leave each calendar year.
• 10 paid holidays per year.
• Up to 160 hours of paid family leave within a rolling 12-month period for eligible employees.
• Short- and long-term disability benefits.
• Life insurance.
• Accidental death and dismemberment insurance.
• Personal accident insurance.
• Critical illness insurance.
• Business travel and accident insurance.
Abacus Group
CTI Staffing
Get handpicked remote jobs straight to your inbox weekly.