
Tech Ops Engineer
Posted Jul 22

Posted Jul 22
This is a fully remote position, open to applicants in United States.
• Oversee and sustain the organization's infrastructure, which includes servers, networks, storage systems, and applications.
• Conduct regular system evaluations and preventive maintenance to guarantee optimal performance and uptime.
• Address system alerts and incidents, diagnosing and resolving issues quickly to minimize downtime.
• Provide technical assistance to resolve infrastructure-related challenges, collaborating closely with other technical teams.
• Diagnose and fix hardware, software, and network issues, escalating to higher-level support when necessary.
• Maintain comprehensive documentation of issues, solutions, and processes to enhance the team's knowledge base.
• Plan and implement system upgrades, patches, and configuration changes, ensuring minimal disruption to business operations.
• Test and validate updates in development environments prior to deployment in production.
• Ensure compliance of all systems with security standards and best practices.
• Identify opportunities for automating routine tasks and processes, enhancing operational efficiency and decreasing manual workload.
• Implement scripts, automation tools, and AI capabilities to streamline system management and monitoring.
• Continuously assess and enhance infrastructure performance, capacity, and resource utilization.
• Support the development and implementation of disaster recovery plans to ensure business continuity in the event of system failures.
• Manage backup and restore processes for critical systems and data, ensuring data integrity and availability.
• Participate in regular disaster recovery testing and drills.
• Plan and execute the decommissioning of legacy infrastructure, coordinating Terraform state cleanup and DNS cutover.
• Collaborate closely with development, network, and security teams to ensure alignment and effective communication on infrastructure projects.
• Provide insights on infrastructure design and architecture to support new projects and initiatives.
• Communicate effectively with non-technical stakeholders, offering updates on system status and issues.
• Bachelor’s degree in Computer Science, Information Technology, or a related field, or equivalent professional experience.
• 5+ years of experience in system administration, or a similar role.
• 5+ years of professional experience in Linux administration, managing AWS resources, developing CI/CD and server orchestration pipelines, scripting, and monitoring.
• Experience with cloud-based production systems at scale.
• Proficient in Amazon Web Services (EC2, VPC, EFS, S3, EKS, etc.).
• Hands-on experience running workloads in Kubernetes (EKS), including ArgoCD GitOps deployments.
• Familiarity with Infrastructure as Code tools, primarily Terraform.
• Experience working within a Python and JavaScript-centric codebase, along with knowledge of their associated best practices.
• Proficient in creating CI/CD pipelines using Jenkins, Concourse, or other CI/CD implementations.
• Experience with monitoring tools, such as Datadog or Prometheus.
• Skilled in scripting for server-side automation, auditing, and monitoring.
• Experience maintaining logging, monitoring, and alerting capabilities using OpenSearch, Vector log pipelines, Prometheus, and Kafka.
• Competent in configuring and managing data sources such as PostgreSQL/Aurora RDS, OpenSearch, Redis, and message streaming platforms like Kafka.
• Design and maintain log ingestion pipelines (e.g., Vector → OpenSearch, Vector → Kafka), including index retention, document shape optimization, and failure recovery.
• Ability to triage and remediate security vulnerabilities (CVEs) across infrastructure components, including container base images, OS packages, and third-party services.
• A working knowledge of modern software practices and technologies, including Agile methodologies.
• Promote and establish development standard methodologies for AWS infrastructure-as-code.
• Familiarity with AWS Well-Architected principles.
• Experience with High Availability implementations.
• Knowledge of Security and Compliance standards.
• Exceptional analytical and problem-solving abilities.
• Intellectual curiosity, a willingness to learn new skills, and the capacity to contribute innovative ideas.
• Health Care Plan (Medical, Dental & Vision)
• Retirement Plan (401k)
• Life Insurance (Basic, Voluntary & AD&D)
• Paid Time Off (Vacation, Sick & Public Holidays)
• Family Leave (Maternity, Paternity)
• Short Term & Long Term Disability
• Training & Development
Global University Systems
GuidePoint Security
Airhosted
Cprime, Inc
Get handpicked remote jobs straight to your inbox weekly.