
Site Reliability Engineer
Posted 6 hours ago

Posted 6 hours ago
This is a fully remote position, open to applicants in Portugal.
• Identify and resolve reliability and performance issues at the application level throughout the pull request process, from nginx to php-fpm, application, database, and cache.
• Optimize and scale the web-serving layer under production load, which includes managing php-fpm pool sizing, process management, nginx timeouts, buffering, upstream keepalive, and capacity planning.
• Set up and troubleshoot nginx as both a reverse proxy and web server, including diagnosing 502/504 error responses.
• Identify and resolve scaling and performance challenges during peak loads.
• Develop, maintain, and provide support for Docker-based workloads.
• Take full ownership of the product deployment and upgrade processes from start to finish.
• Provide support and troubleshoot services that are customer-facing.
• Allocate 50% of working hours collaborating with product teams to define, clarify, implement, and deploy enhancements.
• Collaborate with the team to enhance the product and its underlying infrastructure.
• Establish reliability metrics and ensure they are consistently met across managed systems.
• Engage in on-call and on-duty rotations.
• Work in collaboration with software developers, infrastructure engineers, security personnel, and QA teams on the Shoppingfeed product.
• Minimum of 2 years of proven experience with Linux CLI.
• At least 2 years of demonstrated experience with Docker or similar container technologies.
• Strong practical experience with PHP applications and php-fpm process management, including the ability to read and debug application code.
• Over 5 years of verified experience with Git.
• Familiarity with Linux Kernel, load balancing, PHP, AWS S3, container orchestration, MySQL/MariaDB, and Elasticsearch.
• Proficiency in English at a B2 level.
• Preferred: 3+ years of experience with configuration management tools like Ansible, SaltStack, Puppet, or Chef.
• Preferred: 1+ years of experience with Kubernetes.
• Preferred: 3+ years managing clustered Linux servers.
• Preferred: 2+ years of experience collaborating on large Git projects within a team.
• Preferred: 3+ years using CI/CD tools such as GitHub Actions, GitLab CI, or Jenkins.
• Preferred: Experience with monitoring tools like Datadog, Prometheus/Grafana, or Kibana/Elasticsearch.
• Preferred: 2+ years managing Elasticsearch clusters.
• Preferred: 2+ years managing MySQL/MariaDB databases.
• Proven ability to troubleshoot complex issues effectively.
• Preferred: 3+ years of experience with at least one cloud provider, preferably AWS.
• Preferred: 3+ years of experience with infrastructure-as-code tools, ideally Terraform.
• Experience with RabbitMQ is a plus.
• Diverse role with significant personal responsibility and globally challenging tasks.
• Professional, international work environment characterized by flat hierarchies and quick decision-making processes.
• Supportive and respectful corporate culture that encourages enjoyment at work and innovation.
• Opportunities for learning, development, and growth with backing from the international team.
• Flexible working hours.
• Options for remote or office-based work.
• Competitive salary package.
• Corporate healthcare insurance available after the probationary period.
• Access to the Company Mobility Program.
• Additional paid leave days package, starting at 25 days per year and increasing with tenure.
• Access to therapy sessions through the Employee Assistance Program.
• Modern, well-equipped, and pet-friendly workplace.
Capstone Integrated Solutions
Independence Pet Group
Zocdoc
Fairsource
Get handpicked remote jobs straight to your inbox weekly.