
Data Scientist
Posted 18 hours ago

Posted 18 hours ago
This is a fully remote position, open to applicants in Latin America.
• Develop and manage data pipelines and transformations within Databricks utilizing Python, PySpark, and SQL.
• Conduct exploratory data analysis on marketing, campaign, and customer datasets.
• Create features and prepare datasets for training purposes.
• Train, assess, and optimize machine learning models.
• Assist in the deployment and monitoring of models in production alongside senior engineers.
• Investigate and address data quality issues in source feeds and pipelines.
• Efficiently utilize AI coding tools while ensuring code quality and comprehensibility.
• Engage in sprint planning, daily standups, and retrospectives within an Agile/SCRUM framework.
• Write tests and implement data validation checks.
• Take part in code review processes.
• Document datasets, features, model assumptions, and outcomes.
• Share findings and recommendations with both technical and business stakeholders.
• 1–3 years of practical experience in data science, machine learning, or analytics engineering, including internships, co-ops, research, or complex projects.
• A degree in Computer Science, Statistics, Mathematics, Engineering, Data Science, or equivalent practical experience.
• Strong proficiency in Python, including libraries such as pandas, NumPy, and scikit-learn.
• Solid SQL proficiency, including joins, aggregations, and window functions.
• Understanding of supervised learning, train/test splits, feature engineering, cross-validation, and evaluation metrics such as ROC AUC, RMSE, precision, and recall.
• Working knowledge of descriptive statistics, hypothesis testing, data profiling, and data cleansing.
• Proficient in Git and collaborative development, with familiarity in AI-assisted development tools like Cursor, GitHub Copilot, or Claude.
• Strong verbal and written communication skills in English.
• Experience with Databricks or PySpark is desirable.
• Exposure to MLflow, Delta Lake, or Azure cloud services is a plus.
• Familiarity with time-series forecasting, uplift modeling, or audience segmentation is advantageous.
• Knowledge of Power BI, Tableau, Plotly, or CI/CD pipelines is beneficial.
• Interest in Large Language Models and generative AI applications is a plus.
• U.S. holidays off.
• Generous paid time off (PTO).
• Internal training opportunities.
• Mentorship programs.
• Bi-weekly payment schedule.
• Practical experience with Databricks on a modern lakehouse platform.
• Opportunities for career advancement and code review support from senior architects.
• A collaborative, cross-border culture spanning the U.S. and Latin America.
Amgen
Peraton
Clean Air Task Force
Get handpicked remote jobs straight to your inbox weekly.