
Senior Software Engineer, Data Acquisition
Posted Jul 23

Posted Jul 23
This is a fully remote position, open to applicants in United States.
• Contribute to the design and enhancement of our data acquisition and processing platform to boost reliability, throughput, and observability.
• Utilize and advance web crawling technologies to gather and catalog internet data.
• Construct, manage, and develop large-scale distributed systems that aggregate, process, and deliver data sourced from the web.
• Create and develop backend services that oversee distributed job orchestration, data pipelines, and extensive asynchronous workloads.
• Organize and model the captured data, ensuring high quality and consistency throughout the datasets.
• Continuously enhance the speed, scalability, and fault tolerance of our ingestion systems.
• Collaborate with data product and engineering teams to design and implement new data products driven by the data you help gather, while also improving existing products.
• Acquire and apply domain-specific expertise in web crawling and data acquisition, with guidance from seasoned teammates and access to established systems.
• Over 7 years of professional experience in building or managing backend or infrastructure systems at scale.
• Strong programming skills in Python, Go, Rust, or similar languages, with experience in async/await, coroutines, or concurrency frameworks.
• A solid understanding of software architecture and backend principles; able to clearly articulate concepts related to concurrency, scalability, and fault tolerance.
• Comprehensive knowledge of the browser rendering pipeline and web application architecture (authentication, cookies, HTTP request/response).
• Familiarity with network architecture and debugging techniques (HTTP, DNS, proxies, packet capture, and analysis).
• Deep understanding of distributed systems principles: parallelism, asynchronous programming, backpressure, and message-driven design.
• Experience in designing or maintaining resilient data ingestion, API integration, or ETL systems.
• Proficient in using Linux/Unix command-line tools and managing system resources.
• Familiarity with message queues, orchestration, and distributed task systems (e.g., Kafka, SQS, Airflow).
• Experience in assessing and monitoring data quality, ensuring consistency, completeness, and reliability across releases.
• Stock options.
• Competitive salaries.
• Unlimited paid time off.
• Medical, dental, and vision insurance.
• Health, fitness, and office stipends.
• The permanent option to work from anywhere and in any manner you prefer.
Railroad19
GFT Technologies
Get handpicked remote jobs straight to your inbox weekly.