
Cloud Infrastructure Engineer
Posted 17 hours ago

Posted 17 hours ago
This is a fully remote position, open to applicants in United States.
• Act as the primary point of contact for escalations related to cloud services from customers and application teams.
• Oversee production infrastructure, proactively identifying, isolating, and resolving issues before they affect business operations.
• Execute automation scripts for operations related to cloud resources and operating systems.
• Take ownership of assigned tasks, respond swiftly to production incidents, adhere to troubleshooting and escalation protocols, and ensure service availability and uptime.
• Conduct rapid triage to assess the scope and business impact of cloud infrastructure issues.
• Gather and analyze system logs, network traces, and filesystem data.
• Perform fundamental Red Hat Linux LVM operations, analyze syslogs, and troubleshoot OS-level processes and services.
• Reach out to application owners or customers to obtain missing information.
• Isolate problems across application, operating system, and cloud infrastructure layers.
• Open and manage support cases with AWS and Azure as necessary.
• Prepare logs, evidence, and initial troubleshooting documentation for handoff to the senior L3 Infrastructure team when SLAs are not met.
• Utilize AI assistants and AIOps tools for log analysis, issue detection, summarization, troubleshooting, and evidence collection.
• Monitor AWS and Azure Backup status reports and execute AWS AMI and Azure image-based VM restorations.
• Design and maintain automation scripts in Bash, Python, and Ansible for data collection, filtering, troubleshooting, and evidence gathering.
• Provide support during off-hours, nights, or weekends for critical infrastructure incidents and escalations.
• Must be a U.S. Citizen with no dual citizenship.
• Security Clearance Level Required: Not Applicable.
• 1–3 years of progressive IT experience with hands-on involvement in cloud infrastructure administration and operational support.
• Strong administration skills in Red Hat Linux.
• Practical experience with core services of AWS and Azure.
• Proficiency in programming/scripting languages and automation tools such as Python, Bash, or Ansible.
• Ability to safely execute automation scripts for cloud resources and OS-level tasks, including understanding script inputs and outputs, workflow, and troubleshooting basic execution problems.
• Familiarity with AI tools for troubleshooting, log analysis, and infrastructure support.
• Willingness to actively troubleshoot issues, resolve problems in real-time, and communicate directly with stakeholders.
• Capability to adhere to established procedures, make sound decisions within defined guidelines, and recognize when to escalate issues to senior engineers.
• Eagerness to learn from senior engineers and gradually take on more complex technical responsibilities.
• Experience with AWS and Azure services, including EC2, Virtual Machines, AMI, Managed Images, VPC/VNet, subnets, security groups/NSGs, VPN, Direct Connect, ExpressRoute, load balancers, S3, EBS, EFS, Azure Blobs, Managed Disks, Vault, IAM policies, RBAC, Secret Keys, and backup/restore operations.
• Strong Red Hat Linux administration skills, encompassing LVM/filesystem management, system services and daemons, OS-level logging, performance tuning, and troubleshooting CPU/memory/IO issues.
• 100% paid Medical, Dental & Vision for our employees.
• 6% 401K match (Vested Immediately).
• 29 Days' PTO.
• Flexible Work Schedule.
• Tuition/Certification Reimbursement.
• Growth Opportunities within an Emerging Defense Company.
Chainguard
Montreal Oficial
RunPod
WorkOS
Get handpicked remote jobs straight to your inbox weekly.