
OSS Engineer
Posted Sep 9

Posted Sep 9
This is a fully remote position, open to applicants in California.
• Design, develop, test, deploy, and maintain internal web applications, APIs, command-line tools, and services for use cases related to fault, performance, configuration, inventory, and service assurance.
• Build capabilities surrounding service-provider tenants, cloud services, CPE inventory, topology and dependencies, device reachability, firmware versions, alarms, incidents, and customer impact.
• Implement modular design principles, conduct code reviews, utilize automated testing, manage version control, and oversee CI/CD processes, release notes, rollback procedures, telemetry, and lifecycle ownership.
• Evaluate whether to build, buy, or integrate capabilities, and incorporate both commercial and open-source OSS components.
• Develop role-based dashboards for executives, the NOC, support teams, and service owners.
• Correlate various data sources such as metrics, logs, traces, events, tickets, changes, customer/tenant data, and device telemetry.
• Automate processes including alert enrichment, deduplication, correlation, ticket creation, routing, escalation, stakeholder notification, evidence collection, and operational reporting.
• Integrate with cloud APIs, Kubernetes, observability platforms, ITSM systems, collaboration channels, CI/CD systems, databases, messaging platforms, and customer-facing operational systems.
• Assist with onboarding, provisioning, inventory management, telemetry, command execution, and firmware operations across managed CPE fleets.
• Create analytics for recurring incidents, chronic customers or device cohorts, change-related failures, capacity trends, knowledge gaps, and opportunities for automation.
• Prototype approaches for anomaly detection, correlation, and assisted diagnosis while maintaining explainability and human oversight.
• Operate internal OSS capabilities as production systems with defined availability targets, along with monitoring, alerting, capacity planning, backup, recovery, and documentation of dependencies.
• Implement security measures including authentication, RBAC, least privilege principles, secrets management, certificate handling, audit logging, secure coding practices, and vulnerability remediation.
• Develop runbooks, user guidelines, support models, and ownership records; provide training for NOC and support personnel.
• Offer escalation support for critical tool failures and automation issues, which may involve occasional approved maintenance or incident work outside regular hours.
• Report directly to the Network Operations Manager and collaborate with teams from cloud, DevOps, firmware, QA, Database Engineering, Product, and Security.
• 5+ years of experience in OSS engineering, network automation, NOC tooling, software engineering for operations, or a related field.
• Proven experience in developing OSS capabilities for service providers, broadband operators, managed-network providers, or telecommunications equipment/software vendors.
• Demonstrated knowledge of telecom Operations Support Systems, including fault, performance, configuration, inventory, topology, service assurance, or network/service orchestration.
• Strong programming proficiency in Python, Go, Java, TypeScript/JavaScript, or a similar programming language.
• Experience in delivering maintainable production applications, services, or automation frameworks.
• Skilled in building APIs, event-driven integrations, data pipelines, databases, and user-facing operational dashboards.
• Practical experience with Linux, containers, Kubernetes, and a public cloud platform.
• Experience in integrating observability, ITSM, collaboration, and operational platforms using REST APIs, webhooks, message queues, or event streams.
• Comprehensive understanding of Git, code review processes, automated testing, CI/CD, release management, and rollback procedures.
• Working knowledge of networking concepts including TCP/IP, DNS, DHCP, TLS, routing, APIs, and systematic troubleshooting across distributed systems.
• Ability to collaborate directly with NOC, support, and operations engineers; translate ambiguous operational issues into clear requirements; and assess outcomes.
• Bachelor's degree in computer science, engineering, or equivalent practical experience.
• Experience with cloud-managed CPE, such as broadband gateways, routers, ONTs, Wi-Fi/mesh systems, or similar edge devices.
• Familiarity with TR-069/CWMP, TR-369/USP, TR-181 data models, ACS/USP controllers, device telemetry, and remote lifecycle management.
• Experience with Google Cloud Platform, Kubernetes, Terraform, Helm, and Git-based CI/CD.
• Knowledge of Grafana, Prometheus, OpenTelemetry, Elasticsearch, or similar observability and visualization platforms.
• Experience with Apache Pulsar, Kafka, or other messaging/streaming platforms and relational or NoSQL operational data stores.
• Experience in integrating Jira Service Management or similar ITSM tools and collaboration platforms via APIs and webhooks.
• Familiarity with TM Forum Open APIs, YANG, NETCONF, RESTCONF, gNMI, SNMP, or other telecom/network-management interfaces.
• Experience with workflow engines, low-code orchestration, or event correlation.
• Knowledge of applying RBAC, secrets management, certificate handling, audit logging, and data protection measures.
• Front-end development or UX experience in creating efficient interfaces for operators working under time constraints.
• Must be eligible to work within North America.
• Visa sponsorship is not available.
• Fully remote work opportunity within North America.
• Equal-opportunity and inclusive working environment.
Mercor
RTX
Expel
Qualus
Get handpicked remote jobs straight to your inbox weekly.