
Computer Vision, Applied Research Scientist
Posted 2 days ago

Posted 2 days ago
This is a fully remote position, open to applicants in United States.
• Take ownership of comprehensive experiments on Boon's foundational model, overseeing everything from architecture design to self-supervised pretraining, supervised fine-tuning, and deployment in production.
• Create and assess innovative multi-stage vision architectures aimed at understanding construction drawings.
• Lead decisions regarding backbones, decoders, fusion strategies, loss functions, and training methodologies.
• Conduct thorough experiments incorporating baselines, ablations, and evaluations on authentic construction drawings.
• Explore research avenues to enhance accuracy across various construction trades and scopes.
• Transition models from experimental notebooks to the production inference pipeline.
• Engage hands-on with tools like PyTorch, YOLO, SAM, DINO, and other contemporary computer-vision frameworks.
• Collaborate with ML engineers on deployment strategies, quantization, and serving processes.
• Troubleshoot issues on customer drawings and integrate insights into future training cycles.
• Partner with teams focused on synthetic data, annotation, and infrastructure to address data and compute requirements.
• Collaborate with engineering leadership to shape the accuracy roadmap and strategic direction.
• Draft internal research reports and present findings, trade-offs, and recommendations to engineering leadership.
• Assist in prioritizing data acquisition and annotation efforts.
• Define evaluation datasets and metrics for performance assessment.
• Recognize failure modes and design experiments to tackle them.
• 3–7+ years of experience in computer vision research.
• Proven track record of published research in computer vision or experience deploying production computer vision models at scale.
• Practical expertise in multi-modal dense prediction, encompassing segmentation, detection, or collaborative vision-language tasks.
• Experience in production with modern vision-transformer backbones such as SAM, DINOv2/v3, CLIP, SigLIP, or similar technologies.
• Strong proficiency in PyTorch and experience in training large-scale vision models.
• Capability to transition models from research environments to production inference pipelines.
• Solid understanding of deep-learning principles, including optimization, loss design, regularization, and self-supervised learning methodologies.
• Proficiency in English, both written and verbal.
• Availability to collaborate during California business hours.
• Experience with Graph Neural Networks or relational reasoning architectures is desirable.
• Familiarity with text spotting, OCR, or scene-text detection is advantageous.
• Experience with LoRA, adapters, or parameter-efficient fine-tuning is a plus.
• Background in self-supervised pretraining is beneficial.
• Experience with engineering or technical drawings, document understanding, or layout analysis is a plus.
• Contributions to open-source computer vision research are appreciated.
• Published works in top venues are an added advantage.
• Significant equity in an early-stage, well-funded startup.
• Publication-friendly environment — we encourage our team to publish their work at leading venues.
• Autonomy to propose, advocate for, and conduct your own experiments, delivering successful solutions to real customers.
• Collaborate with a top-tier team from esteemed tech companies and research institutions.
• Work at the crossroads of computer vision research and tangible industry impact within a trillion-dollar market.
• Potential for additional compensation, including discretionary bonuses and/or equity in the company.
Zillow
Analytical Mechanics Associates
Mercor
NVIDIA
Get handpicked remote jobs straight to your inbox weekly.