Improving the gold standard in NCBI GenBank and related databases: DNA sequences from type specimens and type strains

Susanne S. Renner, Mark D. Scherz, Conrad L. Schoch, Marc Gottschling, Miguel Vences

Research output: Contribution to journalArticlepeer-review

1 Scopus citations

Abstract

Scientific names permit humans and search engines to access knowledge about the biodiversity that surrounds us, and names linked to DNA sequences are playing an ever-greater role in search-and-match identification procedures. Here, we analyze how users and curators of the National Center for Biotechnology Information (NCBI) are flagging and curating sequences derived from nomenclatural type material, which is the only way to improve the quality of DNA-based identification in the long run. For prokaryotes, 18,281 genome assemblies from type strains have been curated by NCBI staff and improve the quality of prokaryote naming. For Fungi, type-derived sequences representing over 21,000 species are now essential for fungus naming and identification. For the remaining eukaryotes, however, the numbers of sequences identifiable as type-derived are minuscule, representing only 739 species of arthropods, 1542 vertebrates, and 125 embryophytes. An increase in the production and curation of such sequences will come from (i) sequencing of types or topotypic specimens in museum collections, (ii) the March 2023 rule changes at the International Nucleotide Sequence Database Collaboration requiring more metadata for specimens, and (iii) efforts by data submitters to facilitate curation, including informing NCBI curators about a specimen's type status. We illustrate different type-data submission journeys and provide best-practice examples from a range of organisms. Expanding the number of type-derived sequences in DNA databases, especially of eukaryotes, is crucial for capturing, documenting, and protecting biodiversity.

Original languageEnglish
Pages (from-to)486-494
Number of pages9
JournalSystematic Biology
Volume73
Issue number2
DOIs
StatePublished - Mar 1 2024

Keywords

  • Best-practice examples
  • curation
  • data submission
  • GenBank
  • museomics
  • nomenclatural types
  • taxonomy

Fingerprint

Dive into the research topics of 'Improving the gold standard in NCBI GenBank and related databases: DNA sequences from type specimens and type strains'. Together they form a unique fingerprint.

Cite this