
Site Reliability Engineer
Posted Aug 5

Posted Aug 5
This is a fully remote position, open to applicants in New Jersey.
• Take charge of the daily operations involved in the exchange trading lifecycle, which includes starting up and shutting down the market, enabling and disabling trading, performing pre- and post-session sanity checks, and capturing artifacts related to settlement, clearing, and trade reporting.
• Engage in an on-call rotation for a regulated live marketplace.
• Lead incident response efforts, drive incidents to resolution, write postmortems, and transform one-time fixes into runbooks and automation.
• Operate and enhance the observability stack, which encompasses dashboards, alert quality, service level objectives (SLOs), and time-to-detection metrics.
• Manage and maintain hybrid infrastructure across Kubernetes clusters, cloud accounts, and geographically distributed on-premises data centers.
• Automate infrastructure and operational procedures utilizing Ansible, Terraform, and Jenkins pipelines, while managing secrets with HashiCorp Vault.
• Provide support for the exchange data platform, which includes PostgreSQL, Kafka change-data-capture and streaming pipelines, Redis, as well as backup/restore and disaster recovery procedures.
• Assist with market-maker and partner connectivity, conformance testing, and onboarding of partners.
• Contribute to process enhancements and develop policies and procedures for monitoring, incident management, change control, and exchange operations.
• A minimum of 5 years of experience in a Site Reliability Engineering, DevOps, production engineering, or technical operations role within a 24/7 production environment.
• Strong foundational knowledge of Linux and proficiency in scripting.
• Experience in supporting and troubleshooting Java applications in a production setting, including stack traces, thread dumps, JVM memory management, garbage collection, logs, and metrics.
• Comprehensive understanding of TCP/IP networking, encompassing connections, ports, routing, and firewalls.
• Practical experience operating Kubernetes in a production environment.
• Experience managing infrastructure as code using Terraform and Ansible.
• Familiarity with CI/CD pipelines utilizing Jenkins or similar tools.
• Experience with observability tools such as Datadog, Prometheus, Grafana, or their equivalents.
• Proven track record of being on-call for critical systems.
• Working knowledge of SQL and relational databases, with PostgreSQL preferred.
• Self-motivated individual capable of achieving results with minimal supervision.
• Comfortable working independently as well as collaboratively within a team.
• Excellent communication and organizational skills, particularly in written incident communication and documentation.
• A background or interest in trading, capital markets, exchange operations, or sports betting is advantageous.
• Familiarity with exchange protocols is a plus.
• Previous experience in a regulated industry is a benefit.
• Startup experience is preferred but not mandatory.
• Medical, Dental, and Vision Benefits: The company covers 100% of the employee premium and 50% of premiums for spouses and dependents.
• Short- & Long-Term Disability coverage.
• Group Term Life Insurance and Accidental Death & Dismemberment (AD&D) insurance.
• Voluntary Life Insurance and AD&D options.
• 401(k) retirement plan.
• Equity options available.
• Flexible time off policy.
• MacBooks provided to all employees.
TEKsystems
TEKsystems
Level Data
Level Data
Get handpicked remote jobs straight to your inbox weekly.