Skip to content

Publications and methods

Research

Automated experimentation, computational performance, and agents that learn from experience.

GPU Kernel Scientist · Search agents · Red Dragon AI research · Earlier publications

ES-FoMo III workshop at ICML 2025 · Vancouver, Canada

GPU Kernel Scientist: An LLM-Driven Framework for Iterative Kernel Optimization

Martin Andrews and Sam Witteveen

Summary

An evolutionary framework for improving GPU kernels through automated experiments. LLMs select previous implementations, design optimization experiments, and write new HIP kernels. An external evaluation system checks correctness and provides timing feedback for the next iteration.

How experiments are selected

The system chooses a base kernel and a reference implementation from its growing population. It proposes five experiment plans and selects three distinct candidates: the most innovative, the one with the highest predicted maximum performance benefit, and the one with the highest predicted minimum benefit. This balances novelty and potential improvement with a more conservative estimate of benefit; the estimates guide experiments rather than guarantee results.

Experimental context

The published work targets AMD MI300 hardware in the AMD Developer Challenge 2025. Evaluation provides end-to-end timings for specified input configurations, without access to kernel profiling. Results describe that competition workload and evaluation setup.

The paper linked here was updated after the workshop to include quantitative results. It describes the measured improvements, baselines, and constraints; these are historical experimental results, not a general performance promise.

Citation — GPU Kernel Scientist
@misc{andrews2025gpukernelscientistllmdriven,
  title={GPU Kernel Scientist: An LLM-Driven Framework for Iterative Kernel Optimization},
  author={Martin Andrews and Sam Witteveen},
  year={2025},
  eprint={2506.20807},
  archivePrefix={arXiv},
  primaryClass={cs.LG},
  url={https://arxiv.org/abs/2506.20807}
}

NeurIPS 2025 Workshop on Multi-Turn Interactions in Large Language Models · San Diego, USA

Reinforcement Learning for Long-Horizon Multi-Turn Search Agents

Vivek Kalyan and Martin Andrews

Summary

This work investigates how reinforcement learning can improve an LLM agent’s ability to search for documents over multiple turns. It evaluates a trained 14-billion-parameter model on a benchmark built from Singapore court judgments and explores restrictions on the number of turns during both training and evaluation.

The experiments show the value of learning from search outcomes and allowing longer interaction horizons. Comparisons in the paper apply to its legal document search benchmark and experimental conditions.

Citation — Multi-turn search agents
@misc{kalyan2025multiturnagenticrag,
  title={Reinforcement Learning for Long-Horizon Multi-Turn Search Agents},
  author={Vivek Kalyan and Martin Andrews},
  year={2025},
  eprint={2510.24126},
  archivePrefix={arXiv},
  primaryClass={cs.CL},
  url={https://arxiv.org/abs/2510.24126}
}

Collaborations

Research with Red Dragon AI

Martin co-founded Red Dragon AI with Sam Witteveen. Work with colleagues there spans verifiable reasoning, multi-hop explanations, language generation, efficient models, and visual relationships.

Explore the Red Dragon AI research archive →

Earlier publications

Relationships from Entity Stream

Martin Andrews and Sam Witteveen · ViGIL workshop at NIPS 2017

Abstract

Relational reasoning is a central component of intelligent behavior, but has proven difficult for neural networks to learn. The Relation Network (RN) module was recently proposed by DeepMind to solve such problems, and demonstrated state-of-the-art results on a number of datasets. However, the RN module scales quadratically in the size of the input, since it calculates relationship factors between every patch in the visual field, including those that do not correspond to entities. In this paper, we describe an architecture that enables relationships to be determined from a stream of entities obtained by an attention mechanism over the input field. The model is trained end-to-end, and demonstrates equivalent performance with greater interpretability while requiring only a fraction of the model parameters of the original RN module.

Links

  • Poster version - Presented at the ViGiL workshop at NIPS 2017 in Long Beach, California, USA

  • Full Workshop Paper - the NIPS 2017 paper accepted for the ViGiL workshop

Citation (now also on arXiv)

@misc{andrews2019relationships,
   title={Relationships from Entity Stream},
   author={Martin Andrews and Sam Witteveen},
   year={2019},
   eprint={1909.03315},
   archivePrefix={arXiv},
   primaryClass={cs.CL}
}

Compressing Word Embeddings

Martin Andrews · ICONIP 2016

Abstract

Recent methods for learning vector space representations of words have succeeded in capturing fine-grained semantic and syntactic regularities using large-scale unlabelled text analysis. However, these representations typically consist of dense vectors that require a great deal of storage and cause the internal structure of the vector space to be opaque. A more idealized representation of a vocabulary would be both compact and readily interpretable. With this goal, this paper first shows that Lloyd's algorithm can compress the standard dense vector representation by a factor of 10 without much loss in performance. Then, using that compressed size as a storage budget, we describe a new GPU-friendly factorization procedure to obtain a representation which gains interpretability as a side-effect of being sparse and non-negative in each encoding dimension. Word similarity and word-analogy tests are used to demonstrate the effectiveness of the compressed representations obtained.

Links

Citation

@Inbook{Andrews2016-CompressingWordEmbeddings,
  author="Andrews, Martin",
  editor="Hirose, Akira and Ozawa, Seiichi and Doya, Kenji 
    and Ikeda, Kazushi and Lee, Minho and Liu, Derong",
  title="Compressing Word Embeddings",
  bookTitle="Neural Information Processing: 
    23rd International Conference, ICONIP 2016, Kyoto, Japan, 
    October 16--21, 2016, Proceedings, Part IV",
  year="2016",
  publisher="Springer International Publishing",
  address="Cham",
  pages="413--422",
  isbn="978-3-319-46681-1",
  doi="10.1007/978-3-319-46681-1_50",
  url="http://dx.doi.org/10.1007/978-3-319-46681-1_50"
}

Named Entity Recognition - Training from Experts

Martin Andrews · IES 2015

Abstract

Named Entity Recognition (NER) is a foundational technology for systems designed to process Natural Language documents. However, many existing state-of-the-art systems are difficult to integrate into commercial settings (due their monolithic construction, licensing constraints, or need for corpuses, for example). In this work, a new NER system is described that uses the output of existing systems over large corpuses as its training set, ultimately enabling labelling with (i) better F1 scores; (ii) higher labelling speeds; and (iii) no further dependence on the external software.

Links

  • 'Lite' version - Poster Prize winner at Nvidia ASEAN GPU conference

  • Full Paper - Presented at IES-2015 in Bangkok, Thailand

Citation

@Inbook{Andrews2016-NERfromExperts,
  author="Andrews, Martin",
  editor="Lavangnananda, Kittichai and Phon-Amnuaisuk, Somnuk
    and Engchuan, Worrawat and Chan, Jonathan H.",
  title="Named Entity Recognition Through Learning from Experts",
  bookTitle="Intelligent and Evolutionary Systems: 
    The 19th Asia Pacific Symposium, IES 2015, Bangkok, Thailand, 
    November 2015, Proceedings",
  year="2016",
  publisher="Springer International Publishing",
  address="Cham",
  pages="281--292",
  isbn="978-3-319-27000-5",
  doi="10.1007/978-3-319-27000-5_23",
  url="http://dx.doi.org/10.1007/978-3-319-27000-5_23"
}