
Site Reliability Engineer – SRE
Posted Jul 31

Posted Jul 31
This is a fully remote position, open to applicants in Florida.
• Oversee the health, availability, and performance of Rocket.net servers, services, and customer environments.
• Proactively detect infrastructure problems, performance declines, and potential service interruptions.
• Examine alerts and operational events to ensure platform stability.
• Conduct regular platform health assessments and verify that critical systems are functioning properly.
• Take part in incident response and manage troubleshooting during events affecting customers.
• Convey platform issues, updates, and solutions to appropriate internal teams.
• Offer advanced technical support for VIP clients and those with intricate hosting-related challenges.
• Serve as a senior escalation point for WordPress Support Engineers when deeper technical investigation is needed.
• Diagnose complex issues related to servers, websites, networking, DNS, performance, caching, and hosting infrastructure.
• Assist customers with advanced technical issues that surpass standard WordPress troubleshooting.
• Investigate and resolve problems concerning server resources, application performance, connectivity, and platform behavior.
• Collaborate directly with customers when necessary to provide expert-level technical support.
• Ensure that escalated customer issues are addressed with urgency, ownership, and effective communication.
• Troubleshoot and maintain Linux-based production environments.
• Investigate issues pertaining to NGINX, Apache, PHP-FPM, MySQL/MariaDB, Redis, and other platform services.
• Assist with server maintenance, configuration modifications, and operational enhancements.
• Support security updates, system hardening, and infrastructure best practices.
• Monitor resource utilization and pinpoint capacity or performance issues.
• Contribute to improving monitoring, automation, and operational workflows.
• Collaborate closely with WordPress Support Engineers, Shift Leads, Site Reliability Engineers, and Engineering teams.
• Provide technical guidance and share knowledge with Support teams.
• Aid in creating internal documentation, troubleshooting guides, and knowledge base articles.
• Identify recurring issues and propose improvements to minimize future incidents.
• Engage in incident reviews and root cause analysis.
• Over 3 years of experience in SRE, DevOps, Platform Engineering, or related roles.
• Extensive experience troubleshooting Linux production environments.
• Background in supporting customer-facing technical environments.
• Strong grasp of web hosting technologies such as NGINX, Apache, PHP-FPM, MySQL/MariaDB, and Redis.
• Advanced troubleshooting capabilities across WordPress, servers, DNS, networking, and performance issues.
• Proficient in using the Linux command line (SSH).
• Solid understanding of DNS, HTTP/HTTPS, SSL/TLS, CDN, and caching technologies.
• Familiarity with Cloudflare, WAF, and web performance optimization.
• Experience with monitoring tools and incident response protocols.
• Ability to independently troubleshoot complex issues and clearly articulate technical solutions.
• Exceptional written and verbal communication skills in English.
• Capacity to work effectively under pressure during incidents that impact customers.
• Opportunity to work from anywhere in the world.
• Flexible vacation policy.
• Paid education opportunities.
TEKsystems
TEKsystems
Level Data
Level Data
Get handpicked remote jobs straight to your inbox weekly.