ES-FoMo III workshop at ICML 2025 · Vancouver, Canada
GPU Kernel Scientist: An LLM-Driven Framework for Iterative Kernel Optimization
Martin Andrews and Sam Witteveen
Summary
An evolutionary framework for improving GPU kernels through automated experiments. LLMs select previous implementations, design optimization experiments, and write new HIP kernels. An external evaluation system checks correctness and provides timing feedback for the next iteration.
How experiments are selected
The system chooses a base kernel and a reference implementation from its growing population. It proposes five experiment plans and selects three distinct candidates: the most innovative, the one with the highest predicted maximum performance benefit, and the one with the highest predicted minimum benefit. This balances novelty and potential improvement with a more conservative estimate of benefit; the estimates guide experiments rather than guarantee results.
Experimental context
The published work targets AMD MI300 hardware in the AMD Developer Challenge 2025. Evaluation provides end-to-end timings for specified input configurations, without access to kernel profiling. Results describe that competition workload and evaluation setup.
The paper linked here was updated after the workshop to include quantitative results. It describes the measured improvements, baselines, and constraints; these are historical experimental results, not a general performance promise.
Citation — GPU Kernel Scientist
@misc{andrews2025gpukernelscientistllmdriven,
title={GPU Kernel Scientist: An LLM-Driven Framework for Iterative Kernel Optimization},
author={Martin Andrews and Sam Witteveen},
year={2025},
eprint={2506.20807},
archivePrefix={arXiv},
primaryClass={cs.LG},
url={https://arxiv.org/abs/2506.20807}
}
NeurIPS 2025 Workshop on Multi-Turn Interactions in Large Language Models · San Diego, USA
Reinforcement Learning for Long-Horizon Multi-Turn Search Agents
Vivek Kalyan and Martin Andrews
Summary
This work investigates how reinforcement learning can improve an LLM agent’s ability to search for documents over multiple turns. It evaluates a trained 14-billion-parameter model on a benchmark built from Singapore court judgments and explores restrictions on the number of turns during both training and evaluation.
The experiments show the value of learning from search outcomes and allowing longer interaction horizons. Comparisons in the paper apply to its legal document search benchmark and experimental conditions.
Citation — Multi-turn search agents
@misc{kalyan2025multiturnagenticrag,
title={Reinforcement Learning for Long-Horizon Multi-Turn Search Agents},
author={Vivek Kalyan and Martin Andrews},
year={2025},
eprint={2510.24126},
archivePrefix={arXiv},
primaryClass={cs.CL},
url={https://arxiv.org/abs/2510.24126}
}
Collaborations
Research with Red Dragon AI
Martin co-founded Red Dragon AI with Sam Witteveen. Work with colleagues there spans verifiable reasoning, multi-hop explanations, language generation, efficient models, and visual relationships.
ICML 2025 · Main conference paper
Martin Andrews and Sam Witteveen
A reasoning system that proposes answers and wordplay explanations, then checks formalised reasoning steps with a verifier. Evaluated on the Cryptonite dataset, it produces Python proofs that make the reasoning behind verified answers available for inspection.
LLMs and Cognition workshop at ICML 2024
Martin Andrews and Sam Witteveen
Using LLM-generated Python proofs to check cryptic crossword wordplay and distinguish correct answers from plausible alternatives.
Read the crossword verification paper on arXiv →
TextGraphs-15 workshop at NAACL 2021
Vivek Kalyan, Sam Witteveen, and Martin Andrews
Retrieving and ranking supporting facts for explanations of science questions, using language models trained on expert relevance ratings and ensembles of rankings.
Read the multi-hop explanation paper on arXiv →
TextGraphs-14 workshop at COLING 2020
Yew Ken Chia, Sam Witteveen, and Martin Andrews
Combining recurrent and Transformer layers to model interactions between supporting facts when ranking multi-hop explanations, rather than scoring each question–fact pair independently.
Read the explanation ranking paper on arXiv →
TextGraphs-13 workshop at EMNLP-IJCNLP 2019
Yew Ken Chia, Sam Witteveen, and Martin Andrews
Reconstructing explanations for elementary science questions from supporting statements, using the language of questions and candidate explanations to guide retrieval and ranking.
Read the explanation generation paper on arXiv →
WNGT workshop at EMNLP-IJCNLP 2019
Sam Witteveen and Martin Andrews
Using a language model to rephrase text while working across both individual sentences and longer passages, including paragraphs.
Read the paraphrasing paper on arXiv →
FEVER workshop at EMNLP-IJCNLP 2019
Martin Andrews and Sam Witteveen
Exploring how a smaller language model can answer factual questions using external knowledge, with unsupervised methods that complement its language-model training.
Read the question answering paper on arXiv →
ViGIL workshop at NeurIPS 2018
Martin Andrews, Yew Ken Chia, and Sam Witteveen
Learning a graph of objects and their relationships through attention, with the scene-graph structure read directly from the top layer of a Transformer.
Read the scene graph paper on arXiv →
CDNNRIA workshop at NeurIPS 2018
Yew Ken Chia, Sam Witteveen, and Martin Andrews
Distilling a large Transformer into a smaller convolutional model for text classification when labelled data is scarce, studying the trade-off between model size, inference speed, and accuracy.
Read the model distillation paper on arXiv →
Explore the Red Dragon AI research archive →