返回 X 名人动态

Thomas Wolf:发布基因定位模型 Carbon-A 与跨物种基因候选数据库

中文全文 · AI 翻译

科学家已经对数千个物种的基因组进行了测序。但对于其中大多数物种,没有人知道基因在哪里。

今天,我们发布 Carbon-A,一个可以直接在 DNA 中寻找基因的开放模型。我们用它创建了一个数据库,涵盖从真菌到哺乳动物的 22,617 个物种、5.66 亿个候选基因(我们也会共享这个数据库)。

为什么要做这件事?有几个原因。

  1. 大象很少患癌。部分解释来自它们的基因:一种抑癌基因在人类中只有一份,而大象拥有额外的拷贝。在有人或某种工具找出一个物种的基因之前,你无法针对它研究这样的问题。

  2. 第一种 GLP-1 药物基于希拉毒蜥毒液中的一种肽。没有人会把希拉毒蜥列入优先研究名单。作物的野生近缘种也一样,它们具有抗旱和抗病能力。这些物种中有许多从未进行过基因注释。

  3. 我们对基因的大部分认识来自小鼠、果蝇和酵母等大约十几个模式物种。这种聚焦取得了显著成效,但基因注释流程仍然高度依赖与这些物种的比较,因此容易漏掉其他物种独有的基因,而这些基因可能恰恰最有趣。

Carbon-A 属于一类较新的、直接读取 DNA 的工具。它从已知基因组中学习,但在读取新基因组时,不需要一个近缘物种作为参照。

  1. 即使熟悉的基因组仍有缺漏。我们与 Active Site 和加州大学圣迭戈分校(UCSD)的合作伙伴一起,在猫、鸡、仓鼠和一种实验室植物等常见物种中,发现了 239 个未被参考注释收录的基因的 RNA 证据。

  2. AlphaFold 能预测蛋白质的形状,但前提是已经有人找到了产生该蛋白质的基因。对于没有已注释基因的物种,它就无从着手。

关于安全的几点说明:

Carbon-A 不设计 DNA,也不预测基因的功能。它标记基因在 DNA 中的位置。这也是我们认为开放发布它是正确选择的原因之一。

更多的基因注释也有助于健康研究:数据库覆盖了许多携带或引发疾病的物种,而控制它们传播的疾病的研究,往往从它们的基因入手。

模型、数据库和更多详情,请看 Georgia 的推文串 👇

引用 @cgeorgiaw 的推文
对照原文

Scientists have sequenced the genomes of thousands of species. But for most of them, nobody knows *where the genes are* Today we're releasing Carbon-A, an open model that finds genes directly in DNA. We used it to create a database of 566 million candidate genes across 22,617 species, from fungi to mammals (that we are sharing as well) Why bother? A few reasons 1. Elephants rarely get cancer. Part of the explanation turned up in their genes: extra copies of a tumor-suppressor gene humans have only one of. You can't ask that kind of question about a species until someone, or some tool, has found its genes. 2. The first GLP-1 drug was based on a peptide from Gila monster venom. Nobody would have put the Gila monster on a priority list. Same for wild relatives of crops, which carry resistance to drought and disease. Many of these species have never been annotated. 3. Most of what we know about genes comes from about a dozen model species, like mice, flies and yeast. That focus worked remarkably well, but annotation pipelines still lean heavily on comparisons with them, which makes genes unique to other species easy to miss, and those can be the most interesting ones. Carbon-A belongs to a newer family of tools that read DNA directly. It learned from known genomes, but it doesn't need a close relative to read a new one. 4. Even familiar genomes still have gaps. With our partners at Active Site and UCSD, we found RNA evidence for 239 genes missing from the reference annotations of species as common as cats, chickens, hamsters and a lab plant. 5. AlphaFold can predict the shape of a protein, but only once someone has found the gene that makes it. In a species with no annotated genes, it has nothing to work with. Some notes on safety: Carbon-A doesn't design DNA or predict what a gene does. It marks where genes are in DNA. This is one of the reasons we think releasing it openly is the right call. More annotations also help health research: many species that carry or cause disease are among those covered, and work on controlling the diseases they spread often starts from their genes. Model and database and more details in Georgia's thread 👇

老杨AI实操

微信扫一扫,添加好友

老杨AI实操的微信好友二维码

手机可长按保存图片,再到微信中识别二维码

保存二维码