<p>We develop an open-source package called AnnDictionary to facilitate the parallel, independent analysis of multiple anndata. AnnDictionary is built on top of LangChain and AnnData and supports all common large language model (LLM) providers. AnnDictionary only requires 1 line of code to configure or switch the LLM backend and it contains numerous multithreading optimizations to support the analysis of many anndata and large anndata. We use AnnDictionary to perform the first benchmarking study of all major LLMs at de novo cell-type annotation. LLMs vary greatly in absolute agreement with manual annotation based on model size. Inter-LLM agreement also varies with model size. We find that LLM annotation of most major cell types to be more than 80-90% accurate, and will maintain a leaderboard of LLM cell type annotation. Furthermore, we benchmark these LLMs at functional annotation of gene sets, and find that Claude 3.5 Sonnet recovers close matches of functional gene set annotations in over 80% of test sets.</p>

错误:搜索内容不能为空,请输入英文关键词
错误:关键词超出字数限制,请精简
高级检索

Benchmarking cell type and gene set annotation by large language models with AnnDictionary

  • George Crowley,
  • Robert C. Jones,
  • Mark Krasnow,
  • Angela Oliveira Pisco,
  • Julia Salzman,
  • Nir Yosef,
  • Siyu He,
  • Madhav Mantri,
  • Jessie Aguirre,
  • Ron Garner,
  • Sal Guerrero,
  • William Harper,
  • Resham Irfan,
  • Sophia Mahfouz,
  • Ravi Ponnusamy,
  • Bhavani A. Sanagavarapu,
  • Ahmad Salehi,
  • Ivan Sampson,
  • Chloe Tang,
  • Alan G. Cheng,
  • James M. Gardner,
  • Burnett Kelly,
  • Thurman Slone,
  • Zifa Wang,
  • Anika Choudhury,
  • Sheela Crasta,
  • Chen Dong,
  • Marcus L. Forst,
  • Douglas E. Henze,
  • Jaeyoon Lee,
  • Maurizio Morri,
  • Serena Y. Tan,
  • Sevahn K. Vorperian,
  • Lynn Yang,
  • Marcela Alcántara-Hernádez,
  • Julian Berg,
  • Dhruv Bhatt,
  • Sara Billings,
  • Andrès Gottfried-Blackmore,
  • Jamie Bozeman,
  • Simon Bucher,
  • Elisa Caffrey,
  • Amber Casillas,
  • Rebecca Chen,
  • Matthew Choi,
  • Rebecca N. Culver,
  • Ivana Cvijovic,
  • Ke Ding,
  • Hala Shakib Dhowre,
  • Hua Dong,
  • Kenneth Donaville,
  • Lauren Duan,
  • Xiaochen Fan,
  • Mariko H. Foecke,
  • Francisco X. Galdos,
  • Eliza A. Gaylord,
  • Karen Gonzales,
  • William R. Goodyer,
  • Michelle Griffin,
  • Yuchao Gu,
  • Shuo Han,
  • Jun Yan He,
  • Paul Heinrich,
  • Rebeca Arroyo Hornero,
  • Keliana Hui,
  • Juan C. Irwin,
  • SoRi Jang,
  • Annie Jensen,
  • Saswati Karmakar,
  • Jengmin Kang,
  • Hailey Kang,
  • Soochi Kim,
  • Stewart J. Kim,
  • William Kong,
  • Mallory A. Laboulaye,
  • Daniel Lee,
  • Gyehyun Lee,
  • Elise Lelou,
  • Anping Li,
  • Baoxiang Li,
  • Wan-Jin Lu,
  • Hayley Raquer-McKay,
  • Elvira Mennillo,
  • Lindsay Moore,
  • Elena Montauti,
  • Karim Mrouj,
  • Shravani Mukherjee,
  • Patrick Neuhöfer,
  • Saphia Nguyen,
  • Honor Paine,
  • Jennifer B. Parker,
  • Julia Pham,
  • Kiet T. Phong,
  • Pratima Prabala,
  • Zhen Qi,
  • Joshua Quintanilla,
  • Iulia Rusu,
  • Ali Reza Rais Sadati,
  • Bronwyn Scott,
  • David Seong,
  • Hosu Sin,
  • Hanbing Song,
  • Bikem Soyur,
  • Sean Spencer,
  • Varun R. Subramaniam,
  • Michael Swift,
  • Aditi Swarup,
  • Greg Szot,
  • Aris Taychameekiatchai,
  • Emily Trimm,
  • Stefan Veizades,
  • Sivakamasundari Vijayakumar,
  • Kim Chi Vo,
  • Tian Wang,
  • Timothy Wu,
  • Yinghua Xie,
  • William Yue,
  • Zue Zhang,
  • Angela Detweiler,
  • Honey Mekonen,
  • Norma F. Neff,
  • Sheryl Paul,
  • Amanda Seng,
  • Jia Yan,
  • Deana Rae Crystal Colburg,
  • Balint Laszlo Forgo,
  • Luca Ghita,
  • Frank McCarthy,
  • Aditi Agrawal,
  • Alina Isakova,
  • Kavita Murthy,
  • Alexandra Psaltis,
  • Wenfei Sun,
  • Kyle Awayan,
  • Pierre Boyeau,
  • Robrecht Cannoodt,
  • Leah Dorman,
  • Samuel D’Souza,
  • Can Ergen,
  • Justin Hong,
  • Harper Hua,
  • Erin McGeever,
  • Antoine de Morree,
  • Luise A. Seeker,
  • Alexander J. Tarashansky,
  • Astrid Gillich,
  • Taha A. Jan,
  • Angela Ling,
  • Abhishek Murti,
  • Nikita Sajai,
  • Ryan M. Samuel,
  • Juliane Winkler,
  • Steven E. Artandi,
  • Philip A. Beachy,
  • Mike F. Clarke,
  • Zev Gartner,
  • Linda C. Giudice,
  • Franklin W. Huang,
  • Juliana Idoyaga,
  • Michael G. Kattah,
  • Christin S. Kuo,
  • Diana J. Laird,
  • Michael T. Longaker,
  • Patricia Nguyen,
  • David Y. Oh,
  • Thomas A. Rando,
  • Kristy Red-Horse,
  • Bruce Wang,
  • Albert Y. Wu,
  • Sean M. Wu,
  • Bo Yu,
  • James Zou,
  • Stephen R. Quake

摘要

We develop an open-source package called AnnDictionary to facilitate the parallel, independent analysis of multiple anndata. AnnDictionary is built on top of LangChain and AnnData and supports all common large language model (LLM) providers. AnnDictionary only requires 1 line of code to configure or switch the LLM backend and it contains numerous multithreading optimizations to support the analysis of many anndata and large anndata. We use AnnDictionary to perform the first benchmarking study of all major LLMs at de novo cell-type annotation. LLMs vary greatly in absolute agreement with manual annotation based on model size. Inter-LLM agreement also varies with model size. We find that LLM annotation of most major cell types to be more than 80-90% accurate, and will maintain a leaderboard of LLM cell type annotation. Furthermore, we benchmark these LLMs at functional annotation of gene sets, and find that Claude 3.5 Sonnet recovers close matches of functional gene set annotations in over 80% of test sets.