Remotery

Data Engineer, AI – Distributed Systems

atZignal LabsRemoteUS flagUnited StatesFull-timeData EngineerMid-levelSenior$120k – $140k/year

Posted Aug 6

This is a fully remote position, open to applicants in United States.

📋 Description

• Design and manage batch and streaming pipelines that process and enhance high-volume unstructured data.

• Take full ownership of pipeline components, overseeing everything from implementation to production monitoring.

• Create and expand data pathways that support NLP, LLM, and retrieval services.

• Conduct data preparation, embedding generation, and indexing workflows.

• Integrate and optimize search and vector stores for semantic search, clustering, and real-time retrieval.

• Develop and refine microservices and APIs that provide analytics to enterprise clients.

• Produce clean, well-tested, and maintainable code.

• Engage in code reviews, CI/CD processes, and infrastructure-as-code methodologies.

• Troubleshoot production issues and enhance system reliability.

• Collaborate with Data Science, ML, Product, and Security teams to transition ideas from prototype to production.


⛳️ Requirements

• Minimum of 3 years of experience in building and operating data pipelines in a production environment.

• Strong programming proficiency in a JVM language such as Scala, Java, or Kotlin.

• Competence in Python for data manipulation and scripting tasks.

• Practical experience with a distributed processing framework, particularly Apache Spark.

• Familiarity with a streaming platform, ideally Kafka, including knowledge of consumer groups, offsets, partitioning, and lag management.

• Hands-on experience with AWS and comfort using Docker.

• Capability to operate within a Kubernetes environment.

• Experience with workflow orchestrators like Airflow, Prefect, or Dagster.

• Proficient in SQL and knowledgeable in at least one NoSQL or caching layer, such as Redis, MongoDB, or DynamoDB.

• Strong foundation in computer science principles, particularly in data structures and algorithms.

• Ability to assess performance and correctness in distributed systems.

• Excellent written communication skills and the ability to work asynchronously across U.S. time zones.


🏝️ Benefits

• Fully remote work environment.

• Support for learning Scala for candidates who are proficient in Java or Kotlin.

• Opportunity to collaborate with experienced engineers managing systems at scale.

People also viewed

Railroad1920 hours ago

Senior Data Engineer – GCP, Python, Iceberg, Delta Lake, Kafka, Snowflake, Databricks

US flagUnited States OnlyFull-timeData Engineer$120k – $180k/year
ApplyView job
Livefront20 hours ago

Data Engineer

PE flagPeru OnlyFull-timeData Engineer
ApplyView job
GFT Technologies21 hours ago

Data Engineer, Mid-level

BR flagBrazil OnlyFull-timeData Engineer
ApplyView job
VIDA21 hours ago

Geospatial Data Engineer – Customer & AI Solutions

DE flagGermany OnlyFull-timeData Engineer
ApplyView job
albo22 hours ago

Data Engineer

MX flagMexico OnlyFull-timeData Engineer
ApplyView job
Leega22 hours ago

Engenheiro de Dados Pleno – AWS

BR flagBrazil OnlyFreelanceData Engineer
ApplyView job

Never miss a great job!

Get handpicked remote jobs straight to your inbox weekly.

Trusted by 7,400+ designers