
Embedded AI Engineer, On-Device Models
Posted 4 days ago

Posted 4 days ago
This is a fully remote position, open to applicants in United States.
• Implement Deepgram's Speech and Conversational models on embedded and low-power consumer hardware by establishing the architecture for on-device, real-time inference across various processors and accelerators.
• Enhance models for limited targets utilizing quantization, pruning, distillation, operator fusion, and architecture-specific compilation to adhere to strict latency, memory, power, and thermal constraints.
• Develop and refine performance-critical runtime code (C, C++, and/or Rust) for embedded systems, including bare-metal and real-time operating systems like FreeRTOS and Zephyr.
• Collaborate with industry-standard edge inference runtimes and vendor NPU/DSP toolchains, efficiently mapping model graphs onto on-device accelerators and CPU/GPU/NPU diversity.
• Construct the on-device runtime infrastructure: model packaging, deployment pipelines, over-the-air update systems, and lightweight telemetry for devices operating under limited or intermittent connectivity.
• Create repeatable benchmarking and validation processes across target hardware to measure latency, accuracy, power usage, memory footprint, and resource utilization, while identifying regressions prior to shipment.
• Work alongside silicon and device vendors for SDK integration and performance optimization, ensuring our models run efficiently on new chipsets and reference platforms.
• Partner with Research and Engine teams to guide model architectures towards designs that are friendly to edge deployment from the outset, minimizing optimization challenges at deployment time.
• Proven experience in delivering production systems on resource-constrained hardware, including embedded systems, mobile platforms, edge AI, or compact low-power devices.
• Advanced proficiency in C, C++, and/or Rust, with a background in writing performance-critical code for constrained environments.
• Practical experience in model optimization for on-device application, covering quantization, pruning, knowledge distillation, or architecture-specific compilation.
• Knowledge of edge inference runtimes (e.g., ONNX Runtime, TensorRT, TFLite, ExecuTorch) and/or vendor-specific NPU/DSP toolchains.
• In-depth understanding of hardware-software interactions, including CPU/GPU/NPU/DSP architectures, memory hierarchies, fixed-point/integer arithmetic, and power management, and their impact on inference performance.
• Experience working closely with hardware: in bare-metal or RTOS environments (e.g., FreeRTOS, Zephyr), embedded Linux, or development on microcontrollers and edge SoCs.
• Excellent communication skills and a proactive builder mentality — capable of defining an ambiguous optimization challenge, driving it to measurable outcomes, and clearly articulating the trade-offs involved.
• Offers Equity
• Offers Bonus
• 10% Annual Bonus
Vets2PM
Prime System Solutions
Fairfield Residential
LYNX PROCESS
Get handpicked remote jobs straight to your inbox weekly.