
Senior ML Engineer
Posted Sep 11

Posted Sep 11
This is a fully remote position, open to applicants in Maryland.
• Take full ownership of the recommendation engine from start to finish, encompassing model selection, algorithm design, preprocessing, and implementing safety guardrails.
• Create, assess, and deploy time-series forecasting and statistical models to optimize Kubernetes workloads across CPU, memory, GPU, and JVM heap.
• Develop and uphold the data-quality layer by identifying and filtering anomalies, load-test windows, startup spikes, and autoscaling artifacts from production telemetry.
• Establish and enhance recommendation-quality metrics through regression testing, behavioral validation, and accuracy and safety assessments in production.
• Investigate and address recommendation-quality challenges arising from customer environments by tracing data, preprocessing, and analyzing model behavior.
• Act as the team's authority on machine learning, providing guidance on technical direction for ML inquiries and evaluating model versus heuristic tradeoffs.
• Produce production-grade Python code for models and pipelines, sharing responsibility for message consumption, metrics ingestion, and caching services.
• Prototype and validate new optimization features from research through gradual, feature-flagged deployment.
• Keep up-to-date with time-series forecasting and resource optimization methods, assessing which techniques are beneficial to adopt.
• Master's degree or higher in a quantitative discipline (Computer Science, Machine Learning, Statistics, Applied Mathematics).
• Over 5 years of experience in software engineering.
• A minimum of 3 years in building and managing machine learning or statistical systems in a production environment.
• Proficient in Python at an expert level, including writing typed, tested, production-quality code.
• Strong understanding of numpy or similar array-based numerical computing.
• Practical experience in time-series analysis and forecasting, encompassing seasonality, trend decomposition, anomaly detection, and classical statistical techniques.
• Proven experience in rigorously testing ML systems, including regression testing against established baselines, behavioral validation, and ensuring numerical reproducibility.
• Familiarity with Kubernetes, including knowledge of resource requests and limits, autoscaling behavior, OOM kills, and CPU throttling.
• Ability to take ownership of a production service, which includes managing queues, caches, retries, observability, and troubleshooting customer-environment issues using logs and metrics.
• Excellent written and verbal communication skills.
• Experience with Prophet or comparable forecasting libraries is a plus.
• Familiarity with Prometheus/PromQL and experience managing metrics at scale is advantageous.
• Background in cloud cost optimization, capacity planning, or infrastructure efficiency is beneficial.
• Experience with AWS (S3, Managed Prometheus) is a plus.
• Having served as the ML domain expert on a team of generalists is advantageous.
• Medical/Dental/Vision coverage.
• 401k with Company Match.
• Health & Dependent Care FSA.
• Unlimited PTO.
• 11 Company Holidays.
• Volunteer/Community Engagement Day.
• Tuition Reimbursement.
• Paid Parental Leave.
• Equity Grants.
• Home internet Reimbursement.
Shield AI
Weekday (YC W21)
Roadpass Digital
Get handpicked remote jobs straight to your inbox weekly.