bertil-braun.de

Applied ML Research Engineer · Karlsruhe

Bertil Braun

I build and evaluate end-to-end machine-learning systems across language-model agents, scalable training infrastructure, and reinforcement learning. I work from experiments and model training through distributed systems, rigorous evaluation, and deployment. Most of my work is open source, with several projects available to try live.

Featured research system · Distributed self-play RL · arXiv report

AlphaZero-Style Chess & Go Engine

An end-to-end AlphaZero-style research platform combining native C++ tree search, batched TensorRT inference, and Python-managed replay, distributed training, and evaluation. Trained entirely from self-play, its chess model reached superhuman strength under the Stockfish 13 calibration below on a single consumer-GPU node. The engine also supports Go and powers a live chess application.

3,251estimated Stockfish 13-calibrated Elo at 100,000 searches $43.20training-node rental, excluding development and evaluation 2.5 daysself-play and training on 8× RTX 4070 SUPER
Self-play RL MCTS C++ / CUDA Distributed training Quantization-aware training
White wins in 54 moves against Stockfish 13 at 200,000 nodes. Click to enlarge.

Selected work

Four complete experiment configurations resolve to one shared pretraining artifact, two post-training artifacts, and four evaluation artifacts
LLM infrastructure · Technical report

LLM-Light — Artifact-Addressed Execution

LLM-Light lets researchers submit complete end-to-end experiments while automatically reusing compatible data, training, and evaluation artifacts across runs. Its single-host executor prevents duplicate work, schedules concurrent GPU stages, and preserves completed stages and supported pretraining checkpoints across relaunches.

Technical report Artifact-addressed execution Cross-run artifact reuse Concurrent GPU execution Relaunch recovery
Project details & artifacts →
Candidate outputs are compared in both orders by multiple language-model judges and aggregated into an Elo ranking
LLM evaluation · GEM² 2025

Scalable Automated Evaluation with LLMs

A reference-free framework for ranking open-ended LLM outputs when no ground-truth answer exists. Multiple models judge every pair in both orders, and Elo aggregation produced rankings that agreed strongly with rankings from 20 domain experts.

GEM² 2025 paper Multi-model judging Bidirectional comparisons Pairwise Elo
Method, results & thesis →
Tool-use and turn-taking data feed two focused adapters and a hybrid streaming runtime for the live full-duplex browser agent
Speech · Real-time systems · arXiv report

Voice-Light — Full-Duplex Streaming Voice Agent

A deployed full-duplex voice agent that keeps listening while it speaks, handles interruptions and backchannels, and executes structured tools. The arXiv report covers the complete path from synthetic and human training data through two focused adapters, locked evaluation, and a two-GPU streaming deployment.

Live demo arXiv report Full-duplex streaming Causal turn-taking Synthetic training data
Project details, results & artifacts →
Computer vision · arXiv report · Live demo

GybeLock — Windsurfing Video Intelligence

An offline computer-vision system that turns long-shot windsurfing footage into stable rider-focused videos, preserving identity through camera movement, crossings, and occlusion. In an evaluation using the same saved detections for every tracker, it reconstructed 11 of 21 development videos exactly, versus four for BoT-SORT and one for OC-SORT, while reducing broken-track excess from 845 to 42 against OC-SORT. The end-to-end system is available to try live.

Live demo arXiv report YOLO detection Pose + multi-object tracking
System, evaluation & artifacts →
Graph RL · arXiv report

Graph-Based Traffic Signal Control

A graph-RL interface that scores individual traffic movements and deterministically assembles them into each junction's legal signal phases. One parameter set executed across five heterogeneous city networks and, within the matched grid family, exceeded the recorded max-pressure baseline on an unseen 6×6 network at every tested demand.

arXiv report Graph neural network Variable action spaces PPO · SUMO
Project details →
Layered architecture diagram for the independent multi-agent story narrator
LLM systems · Agents & orchestration

Agentic LLM Systems

Three separate agent-system projects: Temporal-Light, a self-hosted durable workflow engine; an evidence-gated coding runtime built on it; and an independent story narrator built from five A2A agents and four MCP tool servers. Together they demonstrate durable execution, tool-grounded software work, and multi-agent service boundaries.

A2A / MCP Evidence-gated coding Durable execution
Details & three projects →

Also built