David Hall

I build foundation models and the systems needed to train them.

David Hall

I lead Marin, an open lab for building foundation models, at Open Athena. Percy Liang and I started Marin at Stanford CRFM in 2024. We develop models in the open: code, data, experiments, training runs, failures, and results.

I also created Levanter, a JAX framework for large-scale model training. Before returning to research, I co-founded Semantic Machines, a conversational AI startup acquired by Microsoft. Earlier, I did my PhD at UC Berkeley and built Breeze, a numerical computing library for Scala that became part of the foundation of Apache Spark’s MLlib.

Selected work

Marin — An open lab for foundation model research. We train large models from scratch while publishing the code, data, experiments, and training process as we go.

GitHub · current 535B run · writing

Levanter — A legible, scalable JAX framework for training large language models and other foundation models on GPUs and TPUs. Levanter is now developed inside the Marin monorepo; the original repository still has useful historical documentation. Levanter is built on Haliax, the named-tensor library I designed for JAX, which makes large distributed models easier to read and modify.

Breeze — A numerical processing and linear algebra library for Scala. It later became the basis for much of Apache Spark MLlib's linear algebra.

Selected writing

Launching Marin 535B-A23B

Thoughts on Marin's first year at Open Athena, chikungunya, and launching Marin's largest open training run yet.

Open Athena · 2026

Open Development of Frontier AI

Why the process of building frontier AI should be public, not just the models we release.

Open Athena · 2026

Introducing Marin: An Open Lab for Building Foundation Models

What Marin is, why we built it, and how the open-lab model works. Things have gotten messier in the age of agents, but the commitment stands.

Stanford CRFM · 2025

Marin 32B retrospective

A roller-coaster marathon of hardware changes, loss spikes, and mid-run interventions.

Stanford CRFM · 2025

Marin 8B retrospective

Lessons from our first serious attempt to train a competitive open model from scratch.

Stanford CRFM · 2025

A few papers

I don't write many papers these days. These are some favorites.

Studying the History of Ideas Using Topic Models

Hall, Jurafsky, Manning · EMNLP · 2008

Sparser, Better, Faster GPU Parsing

Hall, Berg-Kirkpatrick, Klein · ACL · 2014

Less Grammar, More Features

Hall, Durrett, Klein · ACL · 2014

Ancestry-Constrained Phylogenetic Analysis Supports the Indo-European Steppe Hypothesis

Chang, Cathcart, Hall, Garrett · Language · 2015

Task-Oriented Dialogue as Dataflow Synthesis

Andreas et al. · TACL · 2020

Fantastic Pretraining Optimizers and Where to Find Them

Wen, Hall, Ma, Liang · arXiv · 2025

More on Google Scholar →

Earlier

I co-founded Semantic Machines in 2014. We built a conversational AI system around a dataflow/program-synthesis representation of dialogue; Microsoft acquired the company in 2018, and parts of the technology later shipped in Outlook.

Before that I was in Dan Klein’s group at UC Berkeley, working on natural language processing and computational historical linguistics. Somewhere along the way I also helped build a StarCraft AI that won the first AIIDE StarCraft AI competition.

Elsewhere

GitHub · Google Scholar · X · LinkedIn · Open Athena · Marin · Stanford CRFM

Email: david dot hall at openathena dot ai