Haoling Pu
ML systems engineer · MS in AI & Innovation, Carnegie Mellon '27 · Pittsburgh
I'm a master's student at Carnegie Mellon's School of Computer Science and an ML systems engineer. I like the layer where the model meets the hardware and the product: kernels, serving, and the agents that sit on top.
This summer at Google I taught an agent to keep a team's knowledge base alive. At CMU's Language Technologies Institute I work on speech translation that starts before the sentence ends. Before that, computer science at the University of Michigan.
Things I've built: a sparse attention kernel that made NVIDIA's competition workloads fly, a wiki that writes itself, and a retrieval system with two brains. The full story is in the résumé; the interesting parts are below.
I'm looking for 2027 new-grad ML and software engineering roles. Email is the best way to reach me.

- May2026 - Aug2026
Gave a team's engineering knowledge a brain: an agent that reads the docs, writes the wiki, and keeps it fresh while everyone sleeps.
- Oct2025 - Present
Teaching a model to translate speech before the sentence has ended, without betting wrong on how it ends.
- May2025 - Aug2025
Robotaxi backend. Cold services used to wake up slowly, so I kept a pool of them warm and a request never waits for a container to yawn.
- 2025 – May 2027School of Computer Science · GPA 4.00 / 4.00
- 2021 – May 2025College of Engineering · GPA 3.95 / 4.00

A CUDA kernel that makes DeepSeek-style sparse attention fly on Blackwell. Passed every workload in NVIDIA's competition, dozens of times faster than the reference.
Keyword search and embeddings, fused so each covers the other's blind spots, with a local open-source model doing the talking.
The agent that knows what the team knows. Shipped for real at Google, for teams beyond my own.
- EchoFrame· code private
Hand it a faded old photo, get it back in color half a minute later. Built for my family's albums.
- CMUReference-free data synthesis for simultaneous speech translation· CMU Language Technologies Institute, Li Lab · in progress
Translating speech while it is still being spoken forces early commitments. This work models plausible continuations of the source to decide what is safe to emit now, without needing reference translations.
- Blackwell-optimized sparse attention kernel· NVIDIA MLSys 2026 competition · competition
A CUDA kernel for DeepSeek-style sparse attention on NVIDIA B200. It passed every workload in the competition, dozens of times faster than the PyTorch reference.
- Basketball
- Gym
- Rock climbing
- Music
- Pool