Overview
The WBG Pioneer - AI Data Engineer Intern role involves collecting, cleaning, transforming, and validating data from various sources to support complex platform modernization. The intern will gain hands-on experience in data migration, agile delivery, and data governance.
Key Responsibilities
- Collect data from databases, APIs, cloud storage, files, and other sources.
- Clean and transform raw data by handling missing values, duplicates, incorrect formats, and outliers, and build and maintain ETL/ELT pipelines using Python, SQL, and data-engineering tools.
- Perform data validation and quality checks to ensure accuracy and consistency.
- Work with large datasets using platforms such as Snowflake, Databricks, Spark, or cloud services.
- Create and maintain data tables, schemas, feature datasets, and data documentation.
- Assist with feature engineering for machine-learning models.
- Support vector databases, embeddings, and Retrieval-Augmented Generation pipelines for generative AI applications.
- Collaborate with data engineers, data scientists, ML engineers, and software developers.
- Participate in code reviews, team meetings, testing, and technical documentation.
Required Experience
1–6 years of relevant professional experience.
Qualifications
Currently enrolled in, or in the final year of, an undergraduate or postgraduate (master’s/PhD) program in IT Computer Science, Software Engineering, Information Technology, and Data Science/Analytics.