Overview
The Data Scientist will lead the design and implementation of advanced data analytics frameworks, support the development of data-driven innovation strategies, and build internal capacity for evidence-based decision-making.
Key Responsibilities
- Operationalize and continuously improve the Innovation Analytics methodology – an AI- and ML-enabled approach to identify, classify, and analyze innovation across WBG portfolios. Combine ontology-based concepts, semantic embeddings, LLM-assisted analysis, natural language processing (NLP), and other statistical and machine-learning methods over the structured and unstructured operational and knowledge data.
- Build reproducible data and analytical pipelines for document ingestion, text extraction, sub-component-level innovation detection, lifecycle and maturity classification, replication and diffusion analysis, and portfolio-level aggregation.
- Design and implement rigorous validation and quality-assurance approaches, including expert validation, sampling, error analysis, confidence assessment, comparison across analytical methods, and monitoring of false positives and false negatives.
- Apply the methodology to country, sector, thematic, Trust Fund, IDA, and other WBG use cases, translating management questions into appropriate datasets, analytical approaches, and decision-relevant outputs.
- Develop clear analytical products and visualizations that identify portfolio patterns, replication and scaling opportunities, capability gaps, and other strategic insights for operational teams and senior management.
- Document analytical methods, data transformations, assumptions, validation results, limitations, and methodology changes, and contribute to progressively standardizing and scaling the Innovation Analytics service.
Required Experience
- A minimum of five years of relevant professional experience.
- Strong applied experience in machine learning and natural language processing (NLP), including text classification, semantic embeddings and search, information extraction, or related analysis of large unstructured text collections.
- Demonstrated experience using large language models (LLMs) programmatically for classification, extraction, structured analysis, or similar analytical tasks, including prompt design, systematic validation, and understanding of the limitations of generative AI.
- Strong grounding in quantitative research methods and model evaluation, including sampling and validation, precision/recall trade-offs, error analysis, confidence assessment, and rigorous interpretation of analytical results.
- Demonstrated ability to translate complex management and operational questions into tractable analytical approaches.
- Demonstrated ability to synthesize and visualize findings across sectors, countries, datasets, and analytical methods for technical and non-technical audiences.
- Strong initiative, results orientation, and teamwork skills, with the ability to work effectively across multidisciplinary teams and evolving priorities.
- Experience with user-facing system and analytical product design, interoperability with APIs and data schemas, development finance, MDB operational data, vector databases/RAG and retrieval systems, or LLM agent engineering would be an advantage.
Qualifications
Master’s degree in data science, computer science, statistics, economics, applied mathematics, or a related quantitative field.