Research

Research studies, open-source diagnostics, and engineering case studies in post-training and multimodal systems.

My work combines ML implementation with experiments that test how training and evaluation choices affect the conclusions we draw. Each entry identifies my contribution and the current form of the work.

Reinforcement learning · EvaluationOpen-source software

AgentCreditBench

Exact-oracle diagnostics that separate identifying helpful actions from inducing the intended policy gradient.

My contribution. Built a CPU benchmark for turn-level credit assignment and compared advantage computations across RL frameworks.

Tiny finite-horizon environments make exact reference values available. The suite includes reproducibility scripts and checks against pinned framework implementations.

These are controlled diagnostic environments; agreement does not establish downstream language-model training performance.

LLM post-training · Data selectionResearch study

Probe-guided prompt selection

Investigating whether lightweight probes can select informative prompts before spending compute on full rollouts.

My contribution. Contributed post-training experiments and coding evaluation on OctoThinker-1B, comparing probe-guided selection with GRPO, DAPO, and LILO baselines.

The study uses logistic-regression probes on model activations to estimate prompt usefulness, with predictor updates as the policy changes during training.

Collaborative research at the workshop-manuscript stage; this entry describes my experimental contribution.

Multimodal learning · Distributed systemsEngineering case study

VLM training & FSDP scaling

Implementing a VLM and investigating suspicious scaling measurements across 2, 4, and 8 V100 GPUs.

My contribution. Built the image-token insertion and training pipeline using SigLIP-2, Qwen2.5, and a projector; profiled FSDP communication and tested forward prefetching.

Prefetching changed the profiler trace without improving measured throughput. Activation checkpointing changed the scaling curve, but did not isolate the cause of the original superlinear result.

The write-up documents the throughput metric, repeated measurements, and the limits of the bottleneck interpretation.

Offline reinforcement learning · Diffusion & flow policiesManuscript in preparation

Offline RL for generative policies

Empirical evaluation of offline RL methods using diffusion and flow policies.

My contribution. Designed and conducted multi-seed experiments and ablations across D4RL locomotion and AntMaze, studying gradient clipping, energy temperature, and candidate perturbations.

Investigated optimization instability through training diagnostics and candidate-level analyses.

Ongoing collaborative research on optimization behavior in generative policies.