
Big Data Infrastructure Engineer
Posted Sep 3

Posted Sep 3
This is a fully remote position, open to applicants in United States.
• Utilize AI-assisted tools to enhance troubleshooting, log analysis, scripting, documentation, research, and operational efficiency, ensuring outputs are validated prior to implementation.
• Recognize repetitive operational tasks and create automation solutions using Shell/Bash, Python, Ansible, APIs, or other suitable technologies.
• Assess new AI and automation capabilities, identifying opportunities to enhance infrastructure operations and engineering workflows.
• Manage, configure, troubleshoot, and optimize Linux operating systems in both production and non-production settings.
• Monitor and evaluate CPU, memory, disk, filesystem, network, processes, and system services; perform configuration and performance optimization.
• Oversee PostgreSQL, MySQL/MariaDB, and Redis administration, which includes configuration, access management, backup and recovery, monitoring, troubleshooting, maintenance, and performance optimization.
• Assist with database replication, high availability, backup/recovery, and capacity management.
• Support and maintain Docker and Kubernetes environments, encompassing deployment, configuration, monitoring, troubleshooting, scaling, and cluster administration.
• Assist clustered and distributed platforms with a focus on high availability, replication, failover, load balancing, quorum, capacity management, and disaster recovery.
• Facilitate the installation, configuration, monitoring, administration, and upgrades of Cloudera/Hortonworks and Hadoop-based environments.
• Provide support and troubleshooting for HDFS, YARN, Hive, Spark, HBase, Kafka, Airflow, Superset, and Trino/Presto.
• Conduct production monitoring and support utilizing Zabbix and Grafana; engage in incident management, root cause analysis, and corrective/preventive actions.
• Assist with security integrations and technologies such as Ranger, LDAP, and Kerberos.
• Collaborate with development, infrastructure, and other technical teams on deployments, upgrades, infrastructure changes, troubleshooting, and production support.
• Maintain technical documentation, operational procedures, automation, and infrastructure configuration records.
• 3–5 years of hands-on experience in Linux/System Administration, Database Administration, Big Data Infrastructure, DevOps, or a related infrastructure role.
• Bachelor's Degree in Computer Science, Computer Engineering, Information Technology, or a related field.
• Strong AI-first and automation-focused mindset, with a proven ability to effectively utilize AI-assisted tools in technical workflows and critically validate generated recommendations prior to implementation.
• Proficient in scripting and automation with Shell/Bash; familiarity with Python, Ansible, APIs, or similar technologies is highly preferred.
• Solid hands-on experience in Linux administration, encompassing system configuration, service management, resource management, storage/filesystems, permissions, networking, troubleshooting, and performance enhancement.
• Good practical knowledge of PostgreSQL, MySQL/MariaDB, and Redis administration, including configuration, backup and recovery, user management, monitoring, maintenance, and performance tuning.
• Understanding of database concepts, including connections, transactions, locks, indexing, query performance, replication, and high availability.
• Good hands-on comprehension of Docker and Kubernetes, covering containers, images, pods, deployments, services, storage, networking, monitoring, resource management, and troubleshooting.
• Familiarity with clustering and distributed system concepts, including high availability, replication, failover, load balancing, and quorum.
• Understanding of networking fundamentals such as TCP/IP, DNS, ports, routing, connectivity, and network troubleshooting.
• Familiarity with Big Data concepts and the Hadoop ecosystem, with knowledge or hands-on experience in HDFS, YARN, Hive, Spark, Kafka, and HBase.
• Experience with Cloudera or Hortonworks platforms is highly desirable.
• Knowledge of Zabbix/Grafana, Ranger/LDAP/Kerberos, CI/CD tools, and Trino/Presto is an advantage.
• Excellent troubleshooting, analytical, and problem-solving abilities, with a methodical approach to investigating issues and identifying root causes.
• Capability to work effectively in production environments, collaborate across technical teams, take ownership of assigned tasks, and continuously expand technical knowledge.
• Preferred certifications include RHCSA, RHCE, CKA, PostgreSQL or MySQL-related certifications/training, and Red Hat Ansible or other relevant automation certifications; certifications are a plus but not a substitute for practical experience.
• Remote work arrangement.
• Full-time employment.
Katapult Labs
Magna Legal Services
Huron
Strategic Systems International
Get handpicked remote jobs straight to your inbox weekly.