usearch12_documentation

Documentation for USEARCH v12

https://github.com/rcedgar/usearch12_documentation

Science Score: 31.0%

This score indicates how likely this project is to be science-related based on various indicators:

  • CITATION.cff file
    Found CITATION.cff file
  • codemeta.json file
    Found codemeta.json file
  • .zenodo.json file
  • DOI references
  • Academic publication links
  • Committers with academic emails
  • Institutional organization owner
  • JOSS paper metadata
  • Scientific vocabulary similarity
    Low similarity (0.5%) to scientific vocabulary
Last synced: 11 months ago · JSON representation ·

Repository

Documentation for USEARCH v12

Basic Info
  • Host: GitHub
  • Owner: rcedgar
  • License: gpl-3.0
  • Language: HTML
  • Default Branch: main
  • Size: 617 KB
Statistics
  • Stars: 1
  • Watchers: 2
  • Forks: 2
  • Open Issues: 0
  • Releases: 0
Created about 2 years ago · Last pushed over 1 year ago
Metadata Files
Readme License Citation

README.md

usearch12_manual

Documentation for USEARCH v12

To view as web site https://rcedgar.github.io/usearch12_documentation

Owner

  • Name: Robert Edgar
  • Login: rcedgar
  • Kind: user

Citation (citation.html)

<!DOCTYPE HTML PUBLIC "-//W3C//DTD HTML 4.01//EN" "http://www.w3.org/TR/html4/strict.dtd">
<html lang="en">
 <head>
  <meta content="en-us" http-equiv="Content-Language"/>
  <meta content="text/html; charset=utf-8" http-equiv="Content-Type"/>
  <meta content="no-cache, no-store, must-revalidate" http-equiv="Cache-Control"/>
  <meta content="no-cache" http-equiv="Pragma"/>
  <meta content="0" http-equiv="Expires"/>
  <title>
   allpairs_global command
  </title>
  <link href="stylesx.css" rel="stylesheet" type="text/css"/>
  <style type="text/css">
   body.c4 {background-color:#c0c0c0;}
  div.c3 {position:absolute; top:45px; left:20px; width:830px; background-color:#ffffff; border-width:10px; border-style:solid;border-color:white;}
  span.c2 {font-weight: bold}
  div.c1 {position:absolute; top:10px; left:20px; width:850px; height:60px;}
    .TopButtonPara       { color:white; background-color:rgb(50,100,150); border-color:rgb(50,100,150); font-family:Arial, Helvetica, sans-serif; font-weight:normal; font-size:9pt; text-align:center; border-width:4px; border-style:solid; }
    .TopButton           { color:white; }
    a.TopButton:link     { text-decoration:none; }
    a.TopButton:visited  { text-decoration:none; }
    a.TopButton:hover    { color:orange; }

    .NewButtonPara       { color:white; background-color:rgb(50,100,150); border-color:rgb(50,100,150); font-family:Arial, Helvetica, sans-serif; font-weight:normal; font-size:9pt; text-align:center; border-width:4px; border-style:solid; }
    .NewButton           { color:white; }
    a.NewButton:link     { text-decoration:none; }
    a.NewButton:visited  { text-decoration:none; }
    a.NewButton:hover    { color:orange; }

    .SideButtonPara       { color:white; font-family:Arial, Helvetica, sans-serif; font-size:9pt; font-weight:normal; text-align:center; line-height:18px; }
    .SideButton           { color:white; }
    a.SideButton:link     { text-decoration:none; }
    a.SideButton:visited  { text-decoration:none; }
    a.SideButton:hover    { color:orange; }
  </style>
 </head>
 <body style="background-color:#c0c0c0;">
  <div>
   <a href="https://drive5.com/usearch">
    <img alt="USEARCH v12" src="usearch12_banner.jpg" style="position:absolute; top:40px; left:10px; padding:0px; border:0px;"/>
   </a>
  </div>
  <div style="position:absolute; top:115px; left:10px; width:850px; background-color:#ffffff; min-height:500px">
   <div style="position:relative; float:left; background-color:#696969; width:125px; left: 0px; min-height:500px; padding:5px; height: 125px;">
    <div class="SideButtonPara" style="text-align:center; padding-top:5px;">
     <a class="SideButton" href="index.html">
      Docs home
     </a>
     <br/>
     <hr style="border:0; border-bottom: 1px solid white;"/>
     <a class="SideButton" href="cmds.html">
      Commands
     </a>
     <br/>
     <a class="SideButton" href="topics.html">
      Topics
     </a>
     <br/>
     <a class="SideButton" href="citation.html">
      Publications
     </a>
     <br/>
    </div>
   </div>
   <div class="ManText" style="left:20px; position: absolute; left:135px; width:695px; background-color:white; padding:10px">
    <h1>
     Publications
    </h1>
    <span class="paper_author">
     R.C. Edgar
    </span>
    <span class="paper_year">
     (2018),
    </span>
    <span class="paper_title">
     <span style="font-weight:bold">
      Taxonomy annotation and guide tree errors in 16S rRNA databases
     </span>
    </span>
    ,
    <span class="paper_link">
     <a href="https://doi.org/10.7717/peerj.5030">
      PeerJ 6:e5030
     </a>
    </span>
    <span class="paper_key">
     <br/>
     • Approx. one in five SILVA and Greengenes taxonomy annotations are wrong
    </span>
    <span class="paper_key">
     <br/>
     • SILVA and Greengenes trees have pervasive conflicts with type strain taxonomies
    </span>
    <br/>
    <br/>
    <span class="paper_author">
     R.C. Edgar
    </span>
    <span class="paper_year">
     (2018),
    </span>
    <span class="paper_title">
     <span style="font-weight:bold">
      Accuracy of taxonomy prediction for 16S rRNA and fungal ITS sequences
     </span>
    </span>
    ,
    <span class="paper_link">
     <a href="https://doi.org/10.7717/peerj.4652">
      PeerJ 6:e4652
     </a>
    </span>
    <span class="paper_key">
     <br/>
     • Cross-validation by identity, novel benchmark strategy enabling realistic accuracy estimates
    </span>
    <span class="paper_key">
     <br/>
     • Genus accuracy of best methods is 50% on V4 sequences
    </span>
    <span class="paper_key">
     <br/>
     • Recent algorithms do not improve on RDP Classifier or SINTAX
    </span>
    <br/>
    <br/>
    <span class="paper_author">
     R.C. Edgar and H. Flyvbjerg
    </span>
    <span class="paper_year">
     (2018),
    </span>
    <span class="paper_title">
     <span style="font-weight:bold">
      Octave plots for visualizing diversity of microbial OTUs
     </span>
    </span>
    ,
    <span class="paper_link">
     <a href="https://doi.org/10.1101/389833">
      https://doi.org/10.1101/389833
     </a>
    </span>
    <span class="paper_key">
     <br/>
     • Octave plots visualize alpha diversity as a histogram
    </span>
    <span class="paper_key">
     <br/>
     • Plots show shape and completeness of distribution
    </span>
    <br/>
    <br/>
    <span class="paper_author">
     R.C. Edgar
    </span>
    <span class="paper_year">
     (2018),
    </span>
    <span class="paper_title">
     <span style="font-weight:bold">
      UNCROSS2: identification of cross-talk in 16S rRNA OTU tables
     </span>
    </span>
    ,
    <span class="paper_link">
     <a href="https://doi.org/10.1101/400762">
      https://doi.org/10.1101/400762
     </a>
    </span>
    <span class="paper_key">
     <br/>
     • Cross-talk rate is approx. 1% in many Illumina datasets
    </span>
    <span class="paper_key">
     <br/>
     • Cross-talk can cause false positive core microbiome
    </span>
    <span class="paper_key">
     <br/>
     • UNCROSS2 algorithm for filtering cross-talk
    </span>
    <br/>
    <br/>
    <span class="paper_author">
     R.C. Edgar
    </span>
    <span class="paper_year">
     (2017),
    </span>
    <span class="paper_title">
     <span style="font-weight:bold">
      Accuracy of microbial community diversity estimated by closed- and open-reference OTUs
     </span>
    </span>
    ,
    <span class="paper_link">
     <a href="https://doi.org/10.7717/peerj.3889">
      PeerJ 5:e3889
     </a>
    </span>
    <span class="paper_key">
     <br/>
     • QIIME closed- and open-reference clustering generates huge numbers of spurious OTUs
    </span>
    <span class="paper_key">
     <br/>
     • Closed-reference OTU assignment splits strains and species even when no sequence errors
    </span>
    <span class="paper_key">
     <br/>
     • Closed-reference fails to assign different hyper-variable regions to the same OTU
    </span>
    <span class="paper_key">
     <br/>
     • Closed-reference discards many well-known species that are present in Greengenes
    </span>
    <br/>
    <br/>
    <span class="paper_author">
     R.C. Edgar
    </span>
    <span class="paper_year">
     (2017),
    </span>
    <span class="paper_title">
     <span style="font-weight:bold">
      SEARCH_16S: A new algorithm for identifying 16S ribosomal RNA genes in contigs and chromosomes
     </span>
    </span>
    ,
    <span class="paper_link">
     <a href="https://doi.org/10.1101/124131">
      https://doi.org/10.1101/124131
     </a>
    </span>
    <br/>
    <br/>
    <span class="paper_author">
     R.C. Edgar
    </span>
    <span class="paper_year">
     (2017),
    </span>
    <span class="paper_title">
     <span style="font-weight:bold">
      SINAPS: Prediction of microbial traits from marker gene sequences
     </span>
    </span>
    ,
    <span class="paper_link">
     <a href="https://doi.org/10.1101/124156">
      https://doi.org/10.1101/124156
     </a>
    </span>
    <br/>
    <br/>
    <span class="paper_author">
     R.C. Edgar
    </span>
    <span class="paper_year">
     (2017),
    </span>
    <span class="paper_title">
     <span style="font-weight:bold">
      "UNBIAS: An attempt to correct abundance bias in 16S sequencing, with limited success"
     </span>
    </span>
    ,
    <span class="paper_link">
     <a href="https://doi.org/10.1101/124149">
      https://doi.org/10.1101/124149
     </a>
    </span>
    <span class="paper_key">
     <br/>
     • Read abundance has very low correlation with species abundance
    </span>
    <span class="paper_key">
     <br/>
     • Bias caused by gene copy count variation and primer mismatches
    </span>
    <span class="paper_key">
     <br/>
     • Gene copy count and primer mismatches cannot be accurately predicted
    </span>
    <span class="paper_key">
     <br/>
     • Impossible to correct abundance bias
    </span>
    <br/>
    <br/>
    <span class="paper_author">
     R.C. Edgar
    </span>
    <span class="paper_year">
     (2017),
    </span>
    <span class="paper_title">
     <span style="font-weight:bold">
      Updating the 97% identity threshold for 16S ribosomal RNA OTUs
     </span>
    </span>
    ,
    <span class="paper_link">
     <a href="https://doi.org/10.1093/bioinformatics/bty113">
      Bioinformatics 34(14) 2371-2375
     </a>
    </span>
    <span class="paper_key">
     <br/>
     • Standard 97% OTU identity threshold is too low
    </span>
    <span class="paper_key">
     <br/>
     • Optimal OTU threshold is 99% for full-length 16S, 100% for V4
    </span>
    <br/>
    <br/>
    <span class="paper_author">
     R.C. Edgar
    </span>
    <span class="paper_year">
     (2016),
    </span>
    <span class="paper_title">
     <span style="font-weight:bold">
      UNCROSS: Filtering of high-frequency cross-talk in 16S amplicon reads
     </span>
    </span>
    ,
    <span class="paper_link">
     <a href="https://doi.org/10.1101/088666">
      https://doi.org/10.1101/088666
     </a>
    </span>
    <span class="paper_key">
     <br/>
     • Cross-talk is common, many are reads assigned to wrong sample
    </span>
    <span class="paper_key">
     <br/>
     • UNCROSS algorithm for filtering cross-talk
    </span>
    <br/>
    <br/>
    <span class="paper_author">
     R.C. Edgar
    </span>
    <span class="paper_year">
     (2016),
    </span>
    <span class="paper_title">
     <span style="font-weight:bold">
      UNOISE2: improved error-correction for Illumina 16S and ITS amplicon sequencing
     </span>
    </span>
    ,
    <span class="paper_link">
     <a href="https://doi.org/10.1101/081257">
      https://doi.org/10.1101/081257
     </a>
    </span>
    <span class="paper_key">
     <br/>
     • UNOISE2 algorithm, improved denoiser
    </span>
    <span class="paper_key">
     <br/>
     • Reduces false-positive chimeras compared to UNOISE and DADA2
    </span>
    <br/>
    <br/>
    <span class="paper_author">
     R.C. Edgar
    </span>
    <span class="paper_year">
     (2016),
    </span>
    <span class="paper_title">
     <span style="font-weight:bold">
      UCHIME2: improved chimera prediction for amplicon sequencing
     </span>
    </span>
    ,
    <span class="paper_link">
     <a href="https://doi.org/10.1101/074252">
      https://doi.org/10.1101/074252
     </a>
    </span>
    <span class="paper_key">
     <br/>
     • UCHIME2 algorithm, improved chimera detection
    </span>
    <span class="paper_key">
     <br/>
     • "Fake" chimeras are common, valid biological sequences matching two-parent model
    </span>
    <span class="paper_key">
     <br/>
     • Perfect chimera filtering impossible even with complete and correct reference
    </span>
    <span class="paper_key">
     <br/>
     • Realistic chimera benchmark
    </span>
    <br/>
    <br/>
    <span class="paper_author">
     R.C. Edgar
    </span>
    <span class="paper_year">
     (2016),
    </span>
    <span class="paper_title">
     <span style="font-weight:bold">
      SINTAX: a simple non-Bayesian taxonomy classifier for 16S and ITS sequences
     </span>
    </span>
    ,
    <span class="paper_link">
     <a href="https://doi.org/10.1101/074161">
      https://doi.org/10.1101/074161
     </a>
    </span>
    <span class="paper_key">
     <br/>
     • SINTAX taxonomy prediction algorithm
    </span>
    <span class="paper_key">
     <br/>
     • Fast and simple method, accuracy comparable to RDP Classifier
    </span>
    <br/>
    <br/>
    <span class="paper_author">
     R.C. Edgar and H. Flyvbjerg
    </span>
    <span class="paper_year">
     (2015),
    </span>
    <span class="paper_title">
     <span style="font-weight:bold">
      "Error filtering, pair assembly and error correction for next-generation sequencing reads"
     </span>
    </span>
    ,
    <span class="paper_link">
     <a href="https://doi.org/10.1093/bioinformatics/btv401">
      Bioinformatics 31(21) 3476-3482
     </a>
    </span>
    <span class="paper_key">
     <br/>
     • Quality filtering by expected errors
    </span>
    <span class="paper_key">
     <br/>
     • Bayesian paired read assembler
    </span>
    <span class="paper_key">
     <br/>
     • Most paired read assemblers calculate incorrect Q scores
    </span>
    <span class="paper_key">
     <br/>
     • UNOISE algorithm, first denoiser for Illumina reads
    </span>
    <br/>
    <br/>
    <span class="paper_author">
     R.C. Edgar
     <i>
      et al.
     </i>
    </span>
    <span class="paper_year">
     (2014),
    </span>
    <span class="paper_title">
     <span style="font-weight:bold">
      UCHIME improves sensitivity and speed of chimera detection
     </span>
    </span>
    ,
    <span class="paper_link">
     <a href="https://doi.org/10.1093/bioinformatics/btr381">
      Bioinformatics 27(16) 2194-2200
     </a>
    </span>
    <span class="paper_key">
     <br/>
     • Shows UCHIME faster and more accurate than ChimeraSlayer
    </span>
    <span class="paper_key">
     <br/>
     • This paper report misleading benchmark tests, see critique in UCHIME2 paper
    </span>
    <br/>
    <br/>
    <span class="paper_author">
     R.C. Edgar
    </span>
    <span class="paper_year">
     (2013),
    </span>
    <span class="paper_title">
     <span style="font-weight:bold">
      UPARSE: highly accurate OTU sequences from microbial amplicon reads
     </span>
    </span>
    ,
    <span class="paper_link">
     <a href="https://doi.org/10.1038/nmeth.2604">
      "Nat. Meth.  10, 996-998"
     </a>
    </span>
    <span class="paper_key">
     <br/>
     • Describes UPARSE algorithm for 97% OTU clustering
    </span>
    <span class="paper_key">
     <br/>
     • Stringent error filtering and discarding singletons necessary
    </span>
    <span class="paper_key">
     <br/>
     • Highly accurate OTUs from paired OTUs without full overlap
    </span>
    <br/>
    <br/>
    <span class="paper_author">
     R.C. Edgar
    </span>
    <span class="paper_year">
     (2010),
    </span>
    <span class="paper_title">
     <span style="font-weight:bold">
      Search and clustering orders of magnitude faster than BLAST
     </span>
    </span>
    ,
    <span class="paper_link">
     <a href="https://doi.org/10.1093/bioinformatics/btq461">
      Bioinformatics 26(19) 2460-2461
     </a>
    </span>
    <span class="paper_key">
     <br/>
     • USEARCH algorithm
    </span>
    <span class="paper_key">
     <br/>
     • Default citation for USEARCH software
    </span>
    <br/>
    <br/>
   </div>
  </div>
 </body>
</html>

GitHub Events

Total
  • Issues event: 2
  • Issue comment event: 2
  • Push event: 1
  • Fork event: 1
Last Year
  • Issues event: 2
  • Issue comment event: 2
  • Push event: 1
  • Fork event: 1

Committers

Last synced: about 1 year ago

All Time
  • Total Commits: 14
  • Total Committers: 1
  • Avg Commits per committer: 14.0
  • Development Distribution Score (DDS): 0.0
Past Year
  • Commits: 2
  • Committers: 1
  • Avg Commits per committer: 2.0
  • Development Distribution Score (DDS): 0.0
Top Committers
Name Email Commits
Robert Edgar r****t@d****m 14
Committer Domains (Top 20 + Academic)

Issues and Pull Requests

Last synced: about 1 year ago

All Time
  • Total issues: 1
  • Total pull requests: 1
  • Average time to close issues: about 6 hours
  • Average time to close pull requests: 3 minutes
  • Total issue authors: 1
  • Total pull request authors: 1
  • Average comments per issue: 2.0
  • Average comments per pull request: 1.0
  • Merged pull requests: 0
  • Bot issues: 0
  • Bot pull requests: 0
Past Year
  • Issues: 1
  • Pull requests: 1
  • Average time to close issues: about 6 hours
  • Average time to close pull requests: 3 minutes
  • Issue authors: 1
  • Pull request authors: 1
  • Average comments per issue: 2.0
  • Average comments per pull request: 1.0
  • Merged pull requests: 0
  • Bot issues: 0
  • Bot pull requests: 0
Top Authors
Issue Authors
  • cliffbueno (1)
Pull Request Authors
  • telatin (2)
Top Labels
Issue Labels
Pull Request Labels