
LLM Application Engineer
Posted 1 day ago

Posted 1 day ago
This is a fully remote position, open to applicants in Argentina, +4 more countries.
• Take ownership of production LLM pipelines from start to finish, covering ingestion, multimodal model calls, structured records, storage, confidence flags, retries, and idempotent rescans.
• Develop strategies for prompts and schemas, ensuring schema-aligned or constrained outputs for consistent, product-ready results.
• Create classification and filtering layers, which include taxonomy mapping, demographic or audience filters, deduplication, and cleanup logic.
• Establish and execute evaluation harnesses with golden sets, regression suites, and online metrics.
• Monitor quality metrics for completeness, duplicates, incorrect inclusions, image presence, and related product SLAs.
• Enhance token usage, model tiering, caching, and batching to manage costs and latency effectively.
• Strengthen long-running asynchronous jobs with timeouts, partial recovery, memory limits, and secure production deployments.
• Collaborate with Flutter/mobile and QA teams on field contracts, review queues, and incident debugging.
• Document architecture and runbooks to ensure shared ownership.
• Keep updated on Gemini and peer LLM APIs; suggest model alternatives, fallbacks, and schema enhancements.
• Experience launching LLM applications in production (not just demos): including prompts, structured outputs, retries, observability, and real failure handling.
• Strong hands-on expertise with Google Gemini, particularly in multimodal text, image, and document-style workflows, as well as structured extraction.
• Practical experience with at least one other significant LLM stack, such as OpenAI or Anthropic.
• Ability to perform structured extraction from HTML, PDFs, images, and mixed email-like content.
• Knowledge of classification and taxonomy systems built on LLM outputs.
• Understanding of evaluation discipline: offline evaluations, regression suites, and production quality metrics linked to clear acceptance criteria.
• Awareness of cost and latency considerations: token budgeting, affordable tiers, caching, and batching.
• Experience with Python backend development on serverless cloud platforms like Cloud Functions, and document stores such as Firestore, or similar GCP approaches.
• Proficiency in English at C1 level or higher.
• Availability to overlap with US hours, approximately until 5 PM EST, for live coordination when necessary.
• Demonstrated ownership qualities: providing honest estimates, identifying early blockers, and completing releases.
• Nice to have: experience with schema-aligned LLM frameworks, Google Document AI or other OCR/document intelligence, familiarity with Gmail API/OAuth compliance, computer vision, Firebase/GCP operations, production evaluation corpora, and experience in agency or multi-client studio settings.
• Fully remote and asynchronous-friendly work environment, requiring overlap until approximately 5 PM EST for client or release coordination when necessary.
• Monthly retainer structure with ongoing AI pipeline work for engineers committed to maintaining production quality and cost accountability.
• Genuine autonomy in structuring prompts, schemas, evaluations, and deployments.
• Direct access to PMs, mobile engineers, and key decision-makers.
• Budget for the necessary LLM and cloud tools to expedite your work.
• A peer network of product-minded builders engaged in overlapping client projects.
Planar
AcuityMD
Turquoise Health
CloudPSO
Get handpicked remote jobs straight to your inbox weekly.