
Senior/Principal Local LLM & Generative AI Platform Engineer
Posted 16 hours ago

Posted 16 hours ago
This is a fully remote position, open to applicants in United States.
• Take ownership of the architecture and technical roadmap for a secure, reliable, and maintainable local LLM platform implemented within Parallel Wireless-controlled infrastructure.
• Collaborate with engineering, product, support, IT, information security, legal, and domain experts to prioritize high-value use cases and convert them into actionable product and platform requirements.
• Develop a modular inference and model-gateway layer that incorporates stable APIs, model routing, streaming, concurrency controls, quotas, and interchangeable models or serving backends.
• Assess open-weight language, code, embedding, reranking, and multimodal models against company-specific tasks; document their provenance, licenses, limitations, security posture, hardware requirements, and total cost of ownership.
• Enhance serving capabilities across CPU, GPU, and accelerator resources through continuous batching, caching, parallelism, quantization, and appropriately sized context limits.
• Design and manage RAG and enterprise-search pipelines for approved repositories, wikis, tickets, standards, design documents, test results, logs, and support content.
• Enforce source-system permissions during ingestion and retrieval; integrate identity, SSO, RBAC, secrets management, and audit logging.
• Establish versioned evaluation datasets along with automated offline and online assessments for retrieval quality, groundedness, factual accuracy, citation quality, code correctness, task completion, latency, safety, and refusal behavior.
• Develop release gates and reproducible regression tests; facilitate canary releases, rollbacks, and approval pathways.
• Implement end-to-end observability for model and agent workflows, including traces, errors, latency, throughput, queue time, resource utilization, saturation, availability, and user feedback.
• Design secure tool-calling and agent workflows that maintain least-privilege access, sandboxing, validation, bounded execution, human approval, and traceability.
• Integrate the platform with developer environments, source-control and CI workflows, knowledge systems, ticketing systems, and internal applications via SDKs, APIs, and reference implementations.
• Establish production foundations including CI/CD, configuration and model registries, backups, disaster recovery, capacity planning, vulnerability management, incident response, and lifecycle policies.
• Safeguard proprietary and personal data through network isolation, encryption, retention controls, redaction, secure logging, and defenses against prompt injection, data poisoning, unsafe output handling, and model-supply-chain risks.
• Determine when improvements in prompting or retrieval are sufficient and when fine-tuning, distillation, or other adaptations are warranted.
• Facilitate adoption through documentation, examples, training sessions, office hours, telemetry, and structured feedback.
• Communicate architecture decisions, quality evidence, risks, capacity, and roadmap trade-offs to both technical and business stakeholders.
• BSc or MSc in Computer Science, Computer Engineering, Electrical Engineering, Data Science, or a related field, or equivalent practical experience.
• Typically, 7+ years of hands-on experience in production software, ML platform, search, data, or infrastructure engineering, with significant recent experience in deploying LLM-powered systems; exceptional candidates with equivalent expertise are encouraged to apply.
• Strong Python engineering capabilities and experience in designing maintainable APIs, services, libraries, and data pipelines.
• Deep understanding of transformer-based language models and production inference, including tokenization, context management, batching, KV caching, parallelism, quantization, structured output, tool calling, and common model failure modes.
• Proven experience in building production RAG or enterprise-search systems utilizing embeddings, vector and/or lexical search, metadata filtering, reranking, source attribution, and systematic retrieval evaluation.
• Experience in defining task-specific LLM evaluations using representative datasets, strong baselines, domain-expert review, automated metrics, human feedback, error analysis, and regression thresholds.
• Experience in deploying and managing containerized services on Linux using Docker and Kubernetes or an equivalent orchestration environment.
• Hands-on experience with GPU-backed model serving, performance profiling, capacity planning, monitoring, and reliability engineering.
• Strong knowledge of distributed-system fundamentals, authentication and authorization, API security, secrets management, encryption, auditability, and data lifecycle controls.
• Familiarity with Git, automated testing, CI/CD, infrastructure as code, observability, and production incident response.
• Proficiency in Go, Java, or C/C++ is a plus.
• Experience operating LLMs in on-premises, private-cloud, restricted-network, or air-gapped environments is preferred.
• Familiarity with inference runtimes and serving systems such as vLLM, SGLang, TensorRT-LLM, llama.cpp, Ray Serve, KServe, or Triton is advantageous.
• Experience optimizing inference on NVIDIA and/or AMD GPUs utilizing CUDA, ROCm, profiling tools, tensor parallelism, pipeline parallelism, speculative decoding, prefix/KV caching, or related techniques is beneficial.
• Experience with model and experiment registries, LLM tracing and evaluation platforms, vector databases, hybrid-search engines, and production data-orchestration frameworks is a plus.
• Familiarity with LoRA/QLoRA, dataset curation, synthetic-data generation, distillation, and post-training evaluation is advantageous.
• Proven experience in building code intelligence, repository-aware assistants, developer tools, or IDE and CI integrations for extensive C/C++ and Python codebases is a plus.
• Awareness of Active Directory or another enterprise identity provider, fine-grained document authorization, data-loss prevention, secure software supply chains, model licensing, and AI governance is an advantage.
• Experience in red-teaming LLM or agent systems is a plus.
• Knowledge of telecommunications, 3GPP, RAN/Open RAN, cloud-native network functions, or technical-support workflows is beneficial.
• Contributions to relevant open-source AI, search, MLOps, or infrastructure projects are considered a plus.
• Competitive salary and performance-based bonuses.
• Comprehensive health, dental, and vision insurance.
• Opportunities for professional growth and development.
• Flexible working hours and remote work options.
• Engaging and collaborative work environment.
Intradiem
Coinbase
Astra
Bright Vision Technologies
Get handpicked remote jobs straight to your inbox weekly.