
Site Reliability Engineer
Posted Sep 10

Posted Sep 10
This is a fully remote position, open to applicants in Portugal.
• Oversee and enhance production environments by monitoring availability and adopting a comprehensive approach to system health.
• Develop software, automation, and systems to effectively manage platform infrastructure and applications.
• Elevate reliability, quality, scalability, and speed-to-market for cloud software solutions.
• Assess and improve system performance by identifying capacity constraints, potential failure modes, and performance bottlenecks.
• Deliver operational support and engineering for extensive, distributed applications and services.
• Collect and analyze metrics related to operating systems, applications, networks, and services for performance optimization and fault isolation.
• Collaborate with development teams on thorough testing, release protocols, and practices ensuring production readiness.
• Engage in system design, platform management, capacity planning, incident response, and post-incident enhancements.
• Create sustainable systems and services through automation, improvements in infrastructure, and minimizing operational toil.
• Balance feature delivery with reliability using service level indicators, objectives, and error budgets.
• Enhance the reliability of Voice/UC platforms and integrations, focusing on real-time signaling and media flows, call quality, latency, jitter, packet loss, failover, and overall service availability.
• Collaborate with Voice/UC and AI engineering teams to implement AI-powered communication capabilities, including speech recognition, transcription, summarization, intelligent routing, conversational assistance, and text-to-speech.
• Establish observability for comprehensive Voice/UC and AI service paths.
• Design and test graceful degradation, dependency isolation, retry/fallback strategies, and recovery processes.
• Automate validation and production-readiness assessments for Voice/UC and AI integrations.
• Bachelor's degree in computer science or a related technical or scientific field, or equivalent practical experience.
• 4-7 years of experience in technical roles related to production operations, systems engineering, SRE/DevOps, CI/CD implementation, software deployment, and maintenance of production systems.
• Familiarity with Agile methodologies, DevOps practices, CI/CD pipelines, infrastructure automation, and production monitoring/observability.
• Experience with distributed systems, cloud infrastructure, containers, and dynamic resource management frameworks like Kubernetes.
• Knowledge of distributed storage technologies such as NFS, HDFS, or S3, or similar cloud storage solutions.
• Proficient troubleshooting skills across Linux, applications, networks, APIs, and distributed service dependencies.
• Understanding of Voice/UC or real-time communications concepts, including SIP, RTP/SRTP, WebRTC, SBCs, media services, or similar communication platforms.
• Experience in supporting or integrating AI-enabled services, APIs, or workflows.
• Preferred familiarity with operational aspects of speech/voice AI, machine-learning services, or LLM-based applications.
• Aptitude for utilizing metrics, logs, traces, and service-level indicators to diagnose intricate production issues and drive quantifiable reliability enhancements.
• Proactive mindset in identifying problems, areas for improvement, and performance bottlenecks.
• Strong communication skills across various functions.
• Equal opportunity employer.
• Reasonable accommodations for identified disabilities or limitations as mandated by applicable laws.
• Commitment to diversity and inclusion.
• Opportunities for internal promotion.
• A collaborative and transparent workplace culture.
Horizon3.ai
CLOUD MANTA GmbH
Stefanini LATAM
Akamai Technologies
Get handpicked remote jobs straight to your inbox weekly.