
AI Pipeline Engineer – Security Automation Platform
Posted 2 hours ago

Posted 2 hours ago
This is a fully remote position, open to applicants in Poland, +4 more states.
• Design, build, and manage automated pipelines from start to finish.
• Transform fragile multi-stage batch jobs into resumable, idempotent, and observable systems featuring explicit state machines and recovery paths.
• Establish and uphold latency budgets and SLOs for each stage, ensuring that violations are visible and actionable.
• Create observability layers equipped with metrics, dashboards, alerting, and health gates.
• Design and implement automated holds, rollbacks, blast-radius limits, kill switches, and safe-by-default behaviors.
• Remove manual steps and minimize operational maintenance.
• Develop and maintain unit and integration tests for concurrency, partial failures, external API flakiness, and multi-stage state.
• Investigate and troubleshoot issues across ClickHouse, GitLab CI, S3/object storage, Prometheus/Grafana, and third-party APIs.
• Collaborate with security analysts and the Server team to create architecture and production-ready designs.
• Take ownership of automated protection, progressive release, quality gates, CI at scale, LLM orchestration, and observability for Imunify360's security platform.
• Over 5 years of professional experience in backend, platform, or infrastructure engineering.
• Proven experience in building and operating multi-stage data or automation pipelines.
• Expertise in at least one of the following programming languages: Python, Go, or Rust.
• Strong systems design judgment with experience in creating reliable production systems.
• Hands-on experience with workflow orchestration and job scheduling.
• Knowledge in reliability engineering, encompassing idempotency, retries with backoff, checkpointing, resumability, graceful degradation, backpressure, and partial-failure handling.
• Practical experience with observability tools such as Prometheus/Grafana, LGTM stack, or equivalent, including metrics design.
• Extensive CI/CD experience, preferably with GitLab CI, including dynamic/child pipelines and self-hosted runners.
• Proficient in using Docker and container-based test environments.
• Experience with S3/Ceph or similar object storage solutions.
• Familiarity with ClickHouse or any other columnar database.
• Capability to design state machines and long-running processes that can endure restarts.
• Ability to analyze concurrency across multiple in-flight rollouts.
• Exceptional debugging skills across system, network, and data layers.
• Strong communication skills and the ability to work effectively within a distributed team.
• At least upper-intermediate proficiency in both spoken and written English.
• A strong emphasis on professional development, offering opportunities for learning and growth through engaging and challenging projects, mentorship, and other knowledge-exchange programs.
• Fully remote work with flexible working hours.
• Paid vacation of 24 days per year.
• 10 days of national holidays.
• Unlimited sick leave.
• Compensation for private medical insurance.
• Reimbursement for co-working spaces.
• Gym/sports reimbursement.
• Opportunity to earn a reward for the most innovative idea that the company can patent.
Redpanda Data
Oxfam America
Dragonfli Group
Akamai Technologies
Get handpicked remote jobs straight to your inbox weekly.