
Senior Systems Engineer
Posted 6 hours ago

Posted 6 hours ago
This is a fully remote position, open to applicants in Texas.
• Take initiative to develop, maintain, and support The Home Depot's technical infrastructure, including hardware and system software.
• Design, implement, and provide support for production infrastructure.
• Engage in regular upgrades and application support.
• Perform root cause and post-mortem analyses for security incidents and service disruptions.
• Stay updated on innovations, industry trends, and changes within internal systems.
• Analyze business trends and behavioral data to pinpoint opportunities for improvement and initiatives.
• Assess, develop, and suggest cost-effective technology solutions.
• Research and architect infrastructure, networking, databases, cloud services, AI, and security measures.
• Develop and maintain tools for monitoring and support.
• Contribute to project planning and reporting.
• Work collaboratively with product and project teams to facilitate infrastructure support.
• Assist in the reviews of technology architecture design.
• Oversee and enhance applications, infrastructure, networks, databases, and security.
• Maintain, upgrade, and provide support for systems and infrastructure.
• Resolve vendor problem tickets effectively.
• Create in-house solution documentation.
• Offer application support for production software.
• Keep up-to-date knowledge articles related to infrastructure-as-code and monitoring/alerting.
• Must be at least eighteen years old.
• Must have legal authorization to work in the United States.
• Bachelor’s degree or an equivalent qualification in a relevant field.
• A minimum of 4 years of professional experience is required.
• Ideal candidates should have, or be expected to acquire, extensive technical knowledge in Red Hat OpenShift, Kubernetes, and Cloud Foundry / VMware Tanzu Application Services (TAS).
• Experience with Ansible, ArgoCD, and Flux for platform automation and GitOps deployment processes.
• Proven experience in implementing and optimizing Prometheus, Grafana, and Fluentbit.
• Strong background in high-availability platform architecture, infrastructure-as-code, operational troubleshooting, and enterprise supply chain service resilience.
• On-call rotation is required.
• Experience as a member of a collaborative, cross-functional, modern engineering team.
• Proven troubleshooting and remediation experience across various Information Technology disciplines.
• Experience in installing and upgrading applications or databases, along with system maintenance.
• Familiarity with system and environment analysis, design, and optimization.
• Knowledge of debuggers, runtime analysis, library systems, compiled programming, and software update tools.
• Experience in monitoring, configuring, and tuning systems, networks, or databases.
• Proficiency with operating system commands, utilities, and scripting.
• Experience with cloud platforms like GCP and Azure.
• Experience supporting a 24x7 retail operation.
• Familiarity with version control systems.
• Experience with CI/CD toolchains.
• Knowledge of production system designs, including Infrastructure as Code, High Availability, and Performance monitoring.
• Exposure to Site Reliability Engineering (SRE).
• More than 1 year of prior leadership experience is preferred.
• Flexible remote/virtual work arrangement.
• No travel is necessary.
• Opportunities to develop and enhance technical and leadership skills.
• Access to learning activities, communities of practice, articles, tutorials, and videos for professional growth.
Target Sistemas
PDI Technologies
Palo Alto Networks
Health Care Service Corporation
Get handpicked remote jobs straight to your inbox weekly.