
Intern, Data Engineer
Posted Sep 15

Posted Sep 15
This is a fully remote position, open to applicants in United States.
• Assist in the design, development, and enhancement of contemporary data platforms tailored for analytics, experimentation, and AI-centric workflows.
• Contribute to the construction and upkeep of batch and streaming data pipelines utilizing Spark, Databricks, Snowflake, and cloud-native services.
• Aid in managing ETL/ELT workflows with tools such as Apache Airflow, dbt, or cloud-based schedulers.
• Support the ingestion of structured and semi-structured data from S3, ADLS, GCS, APIs, or Kafka into both raw and curated data layers.
• Develop and sustain SQL and Python transformations aimed at cleaning, joining, and aggregating datasets.
• Engage in data quality evaluations, validation protocols, and basic monitoring tasks.
• Collaborate effectively with data engineers, analysts, data scientists, and AI specialists.
• Prepare datasets and feature tables intended for AI/ML pipelines and autonomous agents.
• Investigate AI-agent interactions with data platforms, encompassing data querying, pipeline triggering, and result summarization.
• Document data flows, schemas, and the logic of pipelines.
• Acquire knowledge and adhere to best practices in data modeling, governance, and privacy.
• Assist with version control and deployment processes using Git alongside fundamental CI/CD workflows.
• Currently enrolled in a Bachelor’s or Master’s program in Computer Science, Data Science, Engineering, Information Systems, or a similar discipline.
• Basic knowledge of SQL, including simple joins, aggregations, and filtering techniques.
• Familiarity with Python for scripting, data manipulation, or as part of coursework projects.
• Introductory understanding of ETL/ELT processes, data lakes, and data warehouse concepts.
• Exposure to at least one cloud service platform: AWS, Azure, or GCP.
• An interest in AI, machine learning, or intelligent systems.
• A strong desire to learn, inquire, and collaborate within a team setting.
• Proficient written and verbal communication skills with a keen attention to detail.
• Expected graduation date between May 2027 and December 2027.
• Preferred: project experience with Databricks, Snowflake, or BigQuery.
• Preferred: exposure to Apache Spark, dbt, or workflow orchestration tools.
• Preferred: familiarity with data formats such as Parquet, JSON, Avro, or Delta Lake.
• Preferred: basic understanding of the differences between streaming and batch processing.
• Preferred: involvement in coursework or projects related to AI agents, LLMs, or ML pipelines.
• Preferred: awareness of data privacy principles such as PII, GDPR, or CCPA.
• Preferred: experience with GitHub or other version control systems.
• Part-time internship schedule: 20–25 hours/week during the semester and up to 40 hours/week during breaks.
Magna Legal Services
Huron
Strategic Systems International
Get handpicked remote jobs straight to your inbox weekly.