Senior/Principal Local LLM & Generative AI Platform Engineer

Posted 16 hours ago

This is a fully remote position, open to applicants in United States.

📋 Description

• Take ownership of the architecture and technical roadmap for a secure, reliable, and maintainable local LLM platform implemented within Parallel Wireless-controlled infrastructure.

• Collaborate with engineering, product, support, IT, information security, legal, and domain experts to prioritize high-value use cases and convert them into actionable product and platform requirements.

• Develop a modular inference and model-gateway layer that incorporates stable APIs, model routing, streaming, concurrency controls, quotas, and interchangeable models or serving backends.

• Assess open-weight language, code, embedding, reranking, and multimodal models against company-specific tasks; document their provenance, licenses, limitations, security posture, hardware requirements, and total cost of ownership.

• Enhance serving capabilities across CPU, GPU, and accelerator resources through continuous batching, caching, parallelism, quantization, and appropriately sized context limits.

• Design and manage RAG and enterprise-search pipelines for approved repositories, wikis, tickets, standards, design documents, test results, logs, and support content.

• Enforce source-system permissions during ingestion and retrieval; integrate identity, SSO, RBAC, secrets management, and audit logging.

• Establish versioned evaluation datasets along with automated offline and online assessments for retrieval quality, groundedness, factual accuracy, citation quality, code correctness, task completion, latency, safety, and refusal behavior.

• Develop release gates and reproducible regression tests; facilitate canary releases, rollbacks, and approval pathways.

• Implement end-to-end observability for model and agent workflows, including traces, errors, latency, throughput, queue time, resource utilization, saturation, availability, and user feedback.

• Design secure tool-calling and agent workflows that maintain least-privilege access, sandboxing, validation, bounded execution, human approval, and traceability.

• Integrate the platform with developer environments, source-control and CI workflows, knowledge systems, ticketing systems, and internal applications via SDKs, APIs, and reference implementations.

• Establish production foundations including CI/CD, configuration and model registries, backups, disaster recovery, capacity planning, vulnerability management, incident response, and lifecycle policies.

• Safeguard proprietary and personal data through network isolation, encryption, retention controls, redaction, secure logging, and defenses against prompt injection, data poisoning, unsafe output handling, and model-supply-chain risks.

• Determine when improvements in prompting or retrieval are sufficient and when fine-tuning, distillation, or other adaptations are warranted.

• Facilitate adoption through documentation, examples, training sessions, office hours, telemetry, and structured feedback.

• Communicate architecture decisions, quality evidence, risks, capacity, and roadmap trade-offs to both technical and business stakeholders.


⛳️ Requirements

• BSc or MSc in Computer Science, Computer Engineering, Electrical Engineering, Data Science, or a related field, or equivalent practical experience.

• Typically, 7+ years of hands-on experience in production software, ML platform, search, data, or infrastructure engineering, with significant recent experience in deploying LLM-powered systems; exceptional candidates with equivalent expertise are encouraged to apply.

• Strong Python engineering capabilities and experience in designing maintainable APIs, services, libraries, and data pipelines.

• Deep understanding of transformer-based language models and production inference, including tokenization, context management, batching, KV caching, parallelism, quantization, structured output, tool calling, and common model failure modes.

• Proven experience in building production RAG or enterprise-search systems utilizing embeddings, vector and/or lexical search, metadata filtering, reranking, source attribution, and systematic retrieval evaluation.

• Experience in defining task-specific LLM evaluations using representative datasets, strong baselines, domain-expert review, automated metrics, human feedback, error analysis, and regression thresholds.

• Experience in deploying and managing containerized services on Linux using Docker and Kubernetes or an equivalent orchestration environment.

• Hands-on experience with GPU-backed model serving, performance profiling, capacity planning, monitoring, and reliability engineering.

• Strong knowledge of distributed-system fundamentals, authentication and authorization, API security, secrets management, encryption, auditability, and data lifecycle controls.

• Familiarity with Git, automated testing, CI/CD, infrastructure as code, observability, and production incident response.

• Proficiency in Go, Java, or C/C++ is a plus.

• Experience operating LLMs in on-premises, private-cloud, restricted-network, or air-gapped environments is preferred.

• Familiarity with inference runtimes and serving systems such as vLLM, SGLang, TensorRT-LLM, llama.cpp, Ray Serve, KServe, or Triton is advantageous.

• Experience optimizing inference on NVIDIA and/or AMD GPUs utilizing CUDA, ROCm, profiling tools, tensor parallelism, pipeline parallelism, speculative decoding, prefix/KV caching, or related techniques is beneficial.

• Experience with model and experiment registries, LLM tracing and evaluation platforms, vector databases, hybrid-search engines, and production data-orchestration frameworks is a plus.

• Familiarity with LoRA/QLoRA, dataset curation, synthetic-data generation, distillation, and post-training evaluation is advantageous.

• Proven experience in building code intelligence, repository-aware assistants, developer tools, or IDE and CI integrations for extensive C/C++ and Python codebases is a plus.

• Awareness of Active Directory or another enterprise identity provider, fine-grained document authorization, data-loss prevention, secure software supply chains, model licensing, and AI governance is an advantage.

• Experience in red-teaming LLM or agent systems is a plus.

• Knowledge of telecommunications, 3GPP, RAN/Open RAN, cloud-native network functions, or technical-support workflows is beneficial.

• Contributions to relevant open-source AI, search, MLOps, or infrastructure projects are considered a plus.


🏝️ Benefits

• Competitive salary and performance-based bonuses.

• Comprehensive health, dental, and vision insurance.

• Opportunities for professional growth and development.

• Flexible working hours and remote work options.

• Engaging and collaborative work environment.

People also viewed

Intradiem10 hours ago

Platform Support Engineer

US flagUnited States OnlyFull-timePlatform Engineer
ApplyView job
Coinbase12 hours ago

Senior Staff Software Engineer, Platform - IAM

US flagUnited States OnlyFull-timePlatform Engineer$253.9k – $298.7k/year
ApplyView job
Astra12 hours ago

Senior Platform Engineer

US flagUnited States OnlyFull-timePlatform Engineer$190k – $230k/year
ApplyView job
Bright Vision Technologies17 hours ago

Virtual Platform Engineer

US flagUnited States OnlyFull-timePlatform Engineer$125k – $145k/year
ApplyView job
Braintrust19 hours ago

Platform Support Engineer

SG flagSingapore OnlyFull-timePlatform Engineer
ApplyView job
YLD20 hours ago

Contract Platform Engineer

GB flagUnited Kingdom OnlyFreelancePlatform Engineer
ApplyView job

Never miss a great job!

Get handpicked remote jobs straight to your inbox weekly.

Trusted by 7,400+ designers