publications

publications by categories in reversed chronological order. generated by jekyll-scholar.

2026

  1. Does RL Expand the Capability Boundary of LLM Agents? A PASS@(k,T) Analysis
    Zhiyuan Zhai, Wenjing Yan, Xiaodan Shao, and 1 more author
    arXiv preprint arXiv:2604.14877, 2026

2025

  1. DeepSeek-R1: Incentivizing Reasoning Capability in LLMs via Reinforcement Learning
    DeepSeek-AI
    arXiv preprint arXiv:2501.12948, 2025
  2. Does Reinforcement Learning Really Incentivize Reasoning Capacity in LLMs Beyond the Base Model?
    Yang Yue, Zhiqi Chen, Rui Lu, and 4 more authors
    arXiv preprint arXiv:2504.13837, 2025
  3. A Sober Look at Progress in Language Model Reasoning: Pitfalls and Paths to Reproducibility
    Andreas Hochlehnert, Hardik Bhatnagar, Vishaal Udandarao, and 3 more authors
    arXiv preprint arXiv:2504.07086, 2025
  4. DAPO: An Open-Source LLM Reinforcement Learning System at Scale
    Qiying Yu, Zheng Zhang, Ruofei Zhu, and 8 more authors
    arXiv preprint arXiv:2503.14476, 2025
  5. Understanding R1-Zero-Like Training: A Critical Perspective
    Zichen Liu, Changyu Chen, Wenjun Li, and 5 more authors
    arXiv preprint arXiv:2503.20783, 2025
  6. Spurious Rewards: Rethinking Training Signals in RLVR
    Rulin Shao, Shuyue Stella Li, Rui Xin, and 8 more authors
    arXiv preprint arXiv:2506.10947, 2025

2024

  1. DeepSeekMath: Pushing the Limits of Mathematical Reasoning in Open Language Models
    Zhihong Shao, Peiyi Wang, Qihao Zhu, and 6 more authors
    arXiv preprint arXiv:2402.03300, 2024
  2. Tulu 3: Pushing Frontiers in Open Language Model Post-Training
    Nathan Lambert, Jacob Morrison, Valentina Pyatkin, and 8 more authors
    arXiv preprint arXiv:2411.15124, 2024
  3. Are Your LLMs Capable of Stable Reasoning?
    Junnan Liu, Hongwei Liu, Linchen Xiao, and 6 more authors
    arXiv preprint arXiv:2412.13147, 2024

2023

  1. Direct Preference Optimization: Your Language Model is Secretly a Reward Model
    Rafael Rafailov, Archit Sharma, Eric Mitchell, and 3 more authors
    Advances in Neural Information Processing Systems (NeurIPS), 2023

2022

  1. Training Language Models to Follow Instructions with Human Feedback
    Long Ouyang, Jeffrey Wu, Xu Jiang, and 8 more authors
    Advances in Neural Information Processing Systems (NeurIPS), 2022
  2. Defining and Characterizing Reward Hacking
    Joar Skalse, Nikolaus H. R. Howe, Dmitrii Krasheninnikov, and 1 more author
    Advances in Neural Information Processing Systems (NeurIPS), 2022

2021

  1. Evaluating Large Language Models Trained on Code
    Mark Chen, Jerry Tworek, Heewoo Jun, and 8 more authors
    arXiv preprint arXiv:2107.03374, 2021
  2. Measuring Mathematical Problem Solving With the MATH Dataset
    Dan Hendrycks, Collin Burns, Saurav Kadavath, and 5 more authors
    NeurIPS Datasets and Benchmarks Track, 2021
  3. Training Verifiers to Solve Math Word Problems
    Karl Cobbe, Vineet Kosaraju, Mohammad Bavarian, and 9 more authors
    arXiv preprint arXiv:2110.14168, 2021

2017

  1. Proximal Policy Optimization Algorithms
    John Schulman, Filip Wolski, Prafulla Dhariwal, and 2 more authors
    arXiv preprint arXiv:1707.06347, 2017

1967

  1. Vision
    wave-mechanics.gif
    Letters on wave mechanics
    Albert Einstein, Erwin Schrödinger, Max Planck, and 2 more authors
    1967

1956

  1. brownian-motion.gif
    Investigations on the Theory of the Brownian Movement
    Albert Einstein
    1956

1950

  1. AJP
    The meaning of relativity
    Albert Einstein and AH Taub
    American Journal of Physics, 1950

1935

  1. Can Quantum-Mechanical Description of Physical Reality Be Considered Complete?
    A. Einstein*†, B. Podolsky*, and N. Rosen*
    Phys. Rev., New Jersey. More Information can be found here , May 1935

1920

  1. Relativity: the Special and General Theory
    Albert Einstein
    1920

1905

  1. Über die von der molekularkinetischen Theorie der Wärme geforderte Bewegung von in ruhenden Flüssigkeiten suspendierten Teilchen
    A. Einstein
    Annalen der physik, 1905
  2. Ann. Phys.
    Un the movement of small particles suspended in statiunary liquids required by the molecular-kinetic theory 0f heat
    A. Einstein
    Ann. Phys., 1905
  3. On the electrodynamics of moving bodies
    A. Einstein
    1905
  4. Ann. Phys.
    Über einen die Erzeugung und Verwandlung des Lichtes betreffenden heuristischen Gesichtspunkt
    Albert Einstein
    Ann. Phys., 1905