
Data Engineer
Posted 11 hours ago

Posted 11 hours ago
This is a fully remote position, open to applicants in United Kingdom, +1 more country.
• Develop and maintain a reliable canonical procurement dataset for analytics and agentic procurement.
• Identify and rectify data quality issues across supplier, agency, solicitation, contract, and awarded-quote data.
• Create and implement efficient, observable, and cost-effective ingestion pipelines from APIs, flat files, scraped pages, PDFs, and various external sources.
• Resolve supplier entities across legal entities, spellings, trading names, subsidiaries, registries, and identifiers.
• Establish shared canonical models for suppliers, contracts, commodities, and agencies.
• Build capabilities for testing, lineage, freshness, provenance, and self-service data.
• Design ingestion and storage strategies for agentic workflows and optimize vector pipelines like Pinecone for RAG.
• Work in collaboration with the Head of Engineering to influence architectural decisions.
• Develop pipelines, handle on-call responsibilities, and investigate raw data sources when discrepancies arise.
• Collaborate with product engineering, customer success, and the Growth Lead.
• Act as the first dedicated data engineering hire and lead the function alongside the Head of Engineering.
• Experience in data engineering within a startup or scale-up environment, focusing on building the platform rather than managing an inherited one.
• Proficiency in SQL and Python, with the insight to determine when a data warehouse is not the appropriate solution to a problem.
• A data model you've designed that has been utilized by others, including aspects you would modify in hindsight.
• Direct experience working with challenging external data sources, such as poorly documented APIs, flat-file drops, scraped pages, PDFs, and uncontrollable data.
• Experience in entity resolution or record linkage, or clear potential for success in that area.
• Commercial awareness to differentiate between data issues that block revenue and those that are merely interesting.
• A proactive approach to diagnosing discrepancies in data before proposing solutions.
• A tendency to create scalable systems rather than one-off pipelines.
• Genuine understanding of AI concepts.
• Resilience and positivity during the fluctuations of a fast-paced startup environment.
• Comfort in operating remotely across time zones with significant autonomy and minimal oversight.
• Curiosity about public-sector data.
• Experience in govtech, public sector, civic tech, public records, or open data is a plus.
• Familiarity with supplier and procurement data, including SAM.gov, UEI, DUNS, NAICS, UNSPSC, cooperative purchasing, COG, or NIGP is advantageous.
• Experience with OpenSearch, Elasticsearch, or vector pipelines such as Pinecone or Milvus integrated with an LLM product is beneficial.
• Bilingual proficiency in English and Spanish is particularly valuable.
• Prior experience as the initial data hire is advantageous.
• Competitive salary and equity opportunities in an early-stage company.
• Comprehensive medical, dental, and vision coverage.
• Flexible paid time off.
• Remote-first work environment, offering genuine flexibility across time zones.
• Full suite of AI tools: Claude Pro, HubSpot, Make, Notion, and more.
• Regular team offsite events, including international gatherings.
• Direct access to the founding team and an opportunity to contribute to impactful projects.
Tonic3
Latitude IT Solutions | SDVOSB
Get handpicked remote jobs straight to your inbox weekly.