
Data Scientist
Posted Aug 4

Posted Aug 4
This is a fully remote position, open to applicants in Poland.
• Train, fine-tune, and optimize LLMs and SLMs utilizing GPU infrastructure.
• Develop and oversee comprehensive machine learning training pipelines.
• Prepare, cleanse, structure, and process substantial amounts of real-world data.
• Choose models, frameworks, tools, and training methodologies based on project needs.
• Implement supervised fine-tuning, transfer learning, prompt tuning, and parameter-efficient fine-tuning.
• Monitor model performance and enhance accuracy, speed, scalability, and resource utilization.
• Work with structured, unstructured, time-series, telemetry, log, and streaming data.
• Document and clarify tools, methods, and technical decisions during the model-training process.
• Collaborate with engineering, data, and business teams to transition models from experimentation to production.
• Troubleshoot issues related to model quality, training stability, GPU performance, and data-pipeline.
• Demonstrated professional experience as a Data Scientist, Machine Learning Engineer, AI Engineer, or in a related role.
• Strong practical experience in training or fine-tuning LLMs and/or SLMs.
• Hands-on experience using GPUs for AI model training.
• Proficient Python programming skills.
• Familiarity with PyTorch, TensorFlow, or Hugging Face Transformers.
• Experience with CUDA, distributed training, cloud GPU platforms, or GPU clusters.
• Thorough understanding of model-training workflows, including data preparation, tokenization, model selection, training, evaluation, and optimization.
• Ability to effectively communicate previous AI projects, tools, technical challenges, and outcomes.
• Experience working with large and complex datasets.
• Strong English communication skills.
• Located in Poland.
• Valuable industry experience in e-commerce, finance or banking, insurance, healthcare or medical data, telemetry and IoT, application or system logs, real-time and streaming data, or high-volume enterprise data environments.
• Nice-to-have experience with distributed model training, LoRA, QLoRA, PEFT, quantization, model compression, MLOps, model deployment, Docker, Kubernetes, MLflow, AWS, Azure, Google Cloud, production AI deployment, and data privacy, security, and governance.
• Participation in the talent pool for 60 months, with the potential for applications to be shared with partners for suitable opportunities.
• Option to withdraw talent-pool consent and have data removed.
Lincoln Institute of Land Policy
General Dynamics Information Technology
CFRA Research
Get handpicked remote jobs straight to your inbox weekly.