
Production Engineer – IC4
Posted 6 days ago

Posted 6 days ago
This is a fully remote position, open to applicants in California, +2 more states.
• Lead the technical transition from legacy systems to contemporary enterprise Linux environments, including upgrade paths from RHEL7 to EL8/EL9.
• Create and manage RPM packages to supplant legacy configurations.
• Oversee the migration of packaging from Chef to CINC.
• Design and implement automated runbooks for consistent migrations.
• Triage, take ownership of, and resolve migration issues from initial report to verified fix.
• Shift monitoring infrastructure to Chronosphere, Prometheus, and Grafana.
• Handle Splunk integrations.
• Integrate services into monitoring and logging stacks, including metrics, dashboards, and alert configurations.
• Strengthen observability frameworks and CI/CD pipelines.
• Develop and sustain reversible rollout and rollback strategies for extensive fleet changes.
• Automate repetitive operational tasks and substitute manual runbook processes with tested and reviewed code.
• Provide Tier-2 operational support and incident response in a follow-the-sun model.
• Conduct hands-on break/fix operations on existing platforms.
• Support application developers with architectural guidance and troubleshooting during cloud migration phases.
• Draft and maintain runbooks, migration strategies, and documentation for package and pipeline ownership.
• Collaborate with the client's SRE organization, internal engineering teams, and customer stakeholders from initial definition to final delivery.
• Over 3 years of professional software engineering experience with a strong focus on hands-on Python.
• Experience in writing and maintaining software utilized by other engineers, encompassing modules, packaging, tests, code reviews, and version control.
• Practical enterprise Linux experience on a fleet scale.
• Proven experience executing large-scale OS upgrades, such as RHEL7 to EL8/EL9 or equivalent major-version migrations.
• Familiarity with building RPM packages to replace legacy configurations.
• Experience with large-scale packaging or configuration management migrations, such as from Chef to CINC.
• Proficiency in triaging, owning, and resolving bugs end to end: reproducing, isolating, fixing, testing, and deploying.
• Capability to enhance CI/CD pipelines, observability frameworks, and rollout/rollback strategies for transitioning from legacy to modern infrastructure.
• Experience in Tier-2 operational support and incident response, including follow-the-sun support and hands-on break/fix operations.
• Experience in onboarding services to newly developed monitoring and logging stacks.
• Proven track record of automating repetitive tasks and documenting technical procedures.
• Must reside in the United States.
• Must possess authorization to work in the United States.
• Preferred: Experience in planning and executing comprehensive logging and monitoring tool rollouts.
• Preferred: Experience supporting cloud cutovers and component migrations to cloud environments.
• Preferred: Experience with Perl scripting.
• Preferred: Familiarity with Chronosphere, Prometheus, Grafana, and Splunk integrations.
• Preferred: Experience with Kickstart/PXE, golden images, or repository and mirror management.
• Certifications should include names, dates, credential IDs, or verification links.
• Full reimbursement of verification costs upon hiring.
• Verified applicants receive priority over non-verified candidates with similar experience.
• Portable verification credential.
Vultr
Canva
Canva
NVIDIA
Get handpicked remote jobs straight to your inbox weekly.