Research
Research studies, open-source diagnostics, and engineering case studies in post-training and multimodal systems.
My work combines ML implementation with experiments that test how training and evaluation choices affect the conclusions we draw. Each entry identifies my contribution and the current form of the work.
AgentCreditBench
Exact-oracle diagnostics that separate identifying helpful actions from inducing the intended policy gradient.
My contribution. Built a CPU benchmark for turn-level credit assignment and compared advantage computations across RL frameworks.
Tiny finite-horizon environments make exact reference values available. The suite includes reproducibility scripts and checks against pinned framework implementations.
These are controlled diagnostic environments; agreement does not establish downstream language-model training performance.
Probe-guided prompt selection
Investigating whether lightweight probes can select informative prompts before spending compute on full rollouts.
My contribution. Contributed post-training experiments and coding evaluation on OctoThinker-1B, comparing probe-guided selection with GRPO, DAPO, and LILO baselines.
The study uses logistic-regression probes on model activations to estimate prompt usefulness, with predictor updates as the policy changes during training.
Collaborative research at the workshop-manuscript stage; this entry describes my experimental contribution.
VLM training & FSDP scaling
Implementing a VLM and investigating suspicious scaling measurements across 2, 4, and 8 V100 GPUs.
My contribution. Built the image-token insertion and training pipeline using SigLIP-2, Qwen2.5, and a projector; profiled FSDP communication and tested forward prefetching.
Prefetching changed the profiler trace without improving measured throughput. Activation checkpointing changed the scaling curve, but did not isolate the cause of the original superlinear result.
The write-up documents the throughput metric, repeated measurements, and the limits of the bottleneck interpretation.
Offline RL for generative policies
Empirical evaluation of offline RL methods using diffusion and flow policies.
My contribution. Designed and conducted multi-seed experiments and ablations across D4RL locomotion and AntMaze, studying gradient clipping, energy temperature, and candidate perturbations.
Investigated optimization instability through training diagnostics and candidate-level analyses.
Ongoing collaborative research on optimization behavior in generative policies.