AI Systems Engineer Intern
Software and Sustainability (S2) Group, Amsterdam
- Deployed and optimized 8 open-source LLMs (0.5B–9B) on-device using llama.cpp, PyTorch, and compression techniques (INT4/INT8 quantization, pruning, distillation), enabling agent-ready NLP inference under strict memory and compute constraints
- Built automated Linux/Bash orchestration pipelines for 480 evaluation runs, collecting ML metrics (latency, energy usage, token throughput, accuracy) to validate autonomous, privacy-preserving language model deployments
- Profiled latency–energy trade-offs across ARM-based mobile hardware and SoC architectures, defining configurations that balance real-time NLP responsiveness with sustainable resource consumption