Intern, Data Engineer

Posted Sep 15

This is a fully remote position, open to applicants in United States.

📋 Description

• Assist in the design, development, and enhancement of contemporary data platforms tailored for analytics, experimentation, and AI-centric workflows.

• Contribute to the construction and upkeep of batch and streaming data pipelines utilizing Spark, Databricks, Snowflake, and cloud-native services.

• Aid in managing ETL/ELT workflows with tools such as Apache Airflow, dbt, or cloud-based schedulers.

• Support the ingestion of structured and semi-structured data from S3, ADLS, GCS, APIs, or Kafka into both raw and curated data layers.

• Develop and sustain SQL and Python transformations aimed at cleaning, joining, and aggregating datasets.

• Engage in data quality evaluations, validation protocols, and basic monitoring tasks.

• Collaborate effectively with data engineers, analysts, data scientists, and AI specialists.

• Prepare datasets and feature tables intended for AI/ML pipelines and autonomous agents.

• Investigate AI-agent interactions with data platforms, encompassing data querying, pipeline triggering, and result summarization.

• Document data flows, schemas, and the logic of pipelines.

• Acquire knowledge and adhere to best practices in data modeling, governance, and privacy.

• Assist with version control and deployment processes using Git alongside fundamental CI/CD workflows.


⛳️ Requirements

• Currently enrolled in a Bachelor’s or Master’s program in Computer Science, Data Science, Engineering, Information Systems, or a similar discipline.

• Basic knowledge of SQL, including simple joins, aggregations, and filtering techniques.

• Familiarity with Python for scripting, data manipulation, or as part of coursework projects.

• Introductory understanding of ETL/ELT processes, data lakes, and data warehouse concepts.

• Exposure to at least one cloud service platform: AWS, Azure, or GCP.

• An interest in AI, machine learning, or intelligent systems.

• A strong desire to learn, inquire, and collaborate within a team setting.

• Proficient written and verbal communication skills with a keen attention to detail.

• Expected graduation date between May 2027 and December 2027.

• Preferred: project experience with Databricks, Snowflake, or BigQuery.

• Preferred: exposure to Apache Spark, dbt, or workflow orchestration tools.

• Preferred: familiarity with data formats such as Parquet, JSON, Avro, or Delta Lake.

• Preferred: basic understanding of the differences between streaming and batch processing.

• Preferred: involvement in coursework or projects related to AI agents, LLMs, or ML pipelines.

• Preferred: awareness of data privacy principles such as PII, GDPR, or CCPA.

• Preferred: experience with GitHub or other version control systems.


🏝️ Benefits

• Part-time internship schedule: 20–25 hours/week during the semester and up to 40 hours/week during breaks.

People also viewed

Katapult Labs1 day ago

AI Data Engineer

CO flagColombia OnlyFull-timeData Engineer
ApplyView job
Magna Legal Services1 day ago

Lead Data Engineer

US flagUnited States OnlyFull-timeData Engineer$155k – $175k/year
ApplyView job
Huron1 day ago

Senior Lead Data Engineer

US flagIllinois OnlyFull-timeData Engineer$140k – $190k/year
ApplyView job
Strategic Systems International1 day ago

Senior AI Data Engineer

MX flagMexico, +1 more countryFull-timeData Engineer
ApplyView job
MGM Resorts International1 day ago

Principal Commercial Data Engineer

US flagFlorida, +7 more statesFull-timeData Engineer$126.2k – $168.3k/year
ApplyView job
Ontrac Solutions1 day ago

Data Architect

US flagMichigan OnlyFreelanceData Engineer$95 – $115/hour
ApplyView job

Never miss a great job!

Get handpicked remote jobs straight to your inbox weekly.

Trusted by 7,400+ designers