References
-
[Tufte 2006] Edward R. Tufte (2006). Beautiful Evidence. Graphics Press, LLC.
-
[Rein 2024] Rein, David and Hou, Betty Li and Stickland, Asa Cooper and Petty, Jackson and Pang, Richard Yuanzhe and Dirani, Julien and Michael, Julian and Bowman, Samuel R (2024). Gpqa: A graduate-level google-proof q&a benchmark. In First Conference on Language Modeling.
-
[Jain 2025] Naman Jain and King Han and Alex Gu and Wen-Ding Li and Fanjia Yan and Tianjun Zhang and Sida Wang and Armando Solar-Lezama and Koushik Sen and Ion Stoica (2025). LiveCodeBench: Holistic and Contamination Free Evaluation of Large Language Models for Code. In The Thirteenth International Conference on Learning Representations. link
-
[Patil 2024] Patil, Shishir G and Zhang, Tianjun and Fang, Vivian and Huang, Roy and Hao, Aaron and Casado, Martin and Gonzalez, Joseph E and Popa, Raluca Ada and Stoica, Ion and others (2024). Goex: Perspectives and designs towards a runtime for autonomous LLM applications. arXiv preprint arXiv:2404.06921.
-
[Parameswaran 2024] Aditya G. Parameswaran and Shreya Shankar and Parth Asawa and Naman Jain and Yujie Wang (2024). Revisiting Prompt Engineering via Declarative Crowdsourcing. In 14th Conference on Innovative Data Systems Research, CIDR 2024, Chaminade, HI, USA, January 14-17, 2024. link
-
[Clavié 2024] Benjamin Clavié (2024). rerankers: A Lightweight Python Library to Unify Ranking Methods. arXiv preprint arXiv:2408.17344. link
-
[Clavié 2025] Benjamin Clavié (2025). ...
Become an O’Reilly member and get unlimited access to this title plus top books and audiobooks from O’Reilly and nearly 200 top publishers, thousands of courses curated by job role, 150+ live events each month,
and much more.
Read now
Unlock full access