
AI Research Engineer – Vision AI, VLM, Physical AI
Posted Sep 2

Posted Sep 2
This is a fully remote position, open to applicants in Washington.
• Develop and enhance models for detection, tracking, segmentation, pose and activity recognition, and scene comprehension.
• Train and assess vision-language models for grounding, dense captioning, temporal question answering, and tool utilization.
• Design retrieval-augmented and agentic loops for tasks involving perception and action.
• Prototype perception-in-the-loop policies utilizing both simulated and real-world data.
• Integrate systems with planners and task graphs for workflows related to manipulation, navigation, and safety.
• Curate datasets, establish evaluation protocols and key performance indicators (KPIs), and conduct ablation studies.
• Package research findings into dependable services using technologies such as Kubernetes, Docker, Ray, and FastAPI.
• Implement profiling, telemetry, and continuous integration (CI) for reproducible scientific outcomes.
• Manage multi-agent pipelines that combine perception, reasoning, simulation, and code generation.
• Convert research from theoretical papers to prototypes and ultimately to deployable production modules.
• Produce publishable or open-source results, or production-ready modules that enhance product KPIs.
• Generate reproducible code, evaluation reports, and capability demonstrations.
• Master’s or Ph.D. in Computer Science, Electrical Engineering, Robotics, or a related field.
• Actively publishing in computer vision, machine learning, or robotics forums such as CVPR, ICCV, ECCV, NeurIPS, ICML, ICLR, CoRL, or RSS.
• Proficiency in PyTorch or JAX and Python programming.
• Experience with CUDA profiling and mixed-precision training.
• Proven research experience in computer vision, with expertise in at least one area such as vision-language models (VLMs), embodied or physical AI, or 3D perception.
• Capability to transition from theoretical papers to coding, conducting ablations, and yielding results with meticulous experiment tracking.
• Preferred: experience with video models, diffusion or 3D generative synthesis/NeRF pipelines, or SLAM/scene reconstruction.
• Preferred: experience in multimodal grounding or temporal reasoning.
• Preferred: familiarity with ROS2, DeepStream/TAO, TensorRT, or ONNX.
• Preferred: experience with Ray and distributed data loaders or sharded checkpoints.
• Preferred: knowledge of testing, linting, profiling, containers, and reproducibility practices.
• Preferred: public GitHub code contributions and first-author publications, or significant open-source impact.
• Make a real impact: research is implemented, driving core features in MVPs and products.
• Receive mentorship from a Principal Architect and senior engineers/researchers.
• Enjoy a balance of leading research methodologies with a pragmatic focus on product development.
Sardine
accesa.eu
Aptura
Cotiviti
Get handpicked remote jobs straight to your inbox weekly.