Loading .travis.yml 0 → 100644 +15 −0 Original line number Diff line number Diff line language: c++ os: linux dist: bionic compiler: gcc before_install: - sudo apt-get install -y valgrind script: - make - export PATH=$PWD/bin:$PATH - git clone https://github.com/frederic-mahe/swarm-tests.git && cd swarm-tests && bash ./run_all_tests.sh | tee tests.log && ! grep -q FAIL tests.log Makefile 0 → 100644 +37 −0 Original line number Diff line number Diff line # SWARM # # Copyright (C) 2012-2019 Torbjorn Rognes and Frederic Mahe # # This program is free software: you can redistribute it and/or modify # it under the terms of the GNU Affero General Public License as # published by the Free Software Foundation, either version 3 of the # License, or (at your option) any later version. # # This program is distributed in the hope that it will be useful, # but WITHOUT ANY WARRANTY; without even the implied warranty of # MERCHANTABILITY or FITNESS FOR A PARTICULAR PURPOSE. See the # GNU Affero General Public License for more details. # # You should have received a copy of the GNU Affero General Public License # along with this program. If not, see <http://www.gnu.org/licenses/>. # # Contact: Torbjorn Rognes <torognes@ifi.uio.no>, # Department of Informatics, University of Oslo, # PO Box 1080 Blindern, NO-0316 Oslo, Norway # Makefile for SWARM PROG=bin/swarm MAN=man/swarm.1 swarm : $(PROG) $(PROG) : make -C src swarm install : $(PROG) $(MAN) /usr/bin/install -c $(PROG) '/usr/local/bin' /usr/bin/install -c $(MAN) '/usr/local/share/man/man1' clean : make -C src clean README.md +37 −33 Original line number Diff line number Diff line [](https://travis-ci.org/torognes/swarm) # swarm A robust and fast clustering method for amplicon-based studies. Loading @@ -16,7 +18,18 @@ To help users, we describe starting from raw fastq files, clustering with **swarm** and producing a filtered OTU table. swarm 2.0 introduces several novelties and improvements over swarm swarm 3.0 introduces: * a much faster default algorithm, * a reduced memory footprint, * binaries for Windows x86-64, GNU/Linux ARM 64, and GNU/Linux POWER8, * an updated, hardened, and thoroughly tested code. Please note that: * strict dereplication of input sequences is now mandatory, * \-\-seeds option (\-w) now outputs results sorted by decreasing abundance, and then by alphabetical order of sequence labels. swarm 2.0 introduced several novelties and improvements over swarm 1.0: * built-in breaking phase now performed automatically, * possibility to output OTU representatives in fasta format (option Loading @@ -24,13 +37,13 @@ swarm 2.0 introduces several novelties and improvements over swarm * fast algorithm now used by default for *d* = 1 (linear time complexity), * a new option called *fastidious* that refines *d* = 1 results and reduces the number of small OTUs, reduces the number of small OTUs. ## Common misconceptions **swarm** is a single-linkage clustering method, with some superficial similarities with other clustering methods (e.g., [Huse et al, 2010](http://www.ncbi.nlm.nih.gov/pmc/articles/PMC2909393/)). **swarm**'s similarities with other clustering methods (e.g., [Huse et al, 2010](http://www.ncbi.nlm.nih.gov/pmc/articles/PMC2909393/)). **swarm**'s novelty is its iterative growth process and the use of sequence abundance values to delineate OTUs. **swarm** properly delineates large OTUs (high recall), and can distinguish OTUs with as little as Loading Loading @@ -76,18 +89,18 @@ cgtcgtcgtcgtcgt where sequence identifiers are unique and end with a value indicating the number of occurrences of the sequence (e.g., `_1000`). Alternative format is possible with the option `-z`, please see the [user manual](https://github.com/torognes/swarm/blob/master/man/swarm_manual.pdf). Swarm format is possible with the option `-z`, please see the [user manual](https://github.com/torognes/swarm/blob/master/man/swarm_manual.pdf). Swarm **requires** each fasta entry to present a number of occurrences to work properly. That crucial information can be produced during the [dereplication](#dereplication) step. [dereplication](#dereplication-mandatory) step. Use `swarm -h` to get a short help, or see the [user manual](https://github.com/torognes/swarm/blob/master/man/swarm_manual.pdf) for a complete description of input/output formats and command line options. The memory footprint of **swarm** is roughly 1.6 times the size of the The memory footprint of **swarm** is roughly 0.6 times the size of the input fasta file. When using the fastidious option, memory footprint can increase significantly. See options `-c` and `-y` to control and cap swarm's memory consumption. Loading @@ -105,14 +118,14 @@ using the ```sh git clone https://github.com/torognes/swarm.git cd swarm/src/ cd swarm/ make cd ../bin/ ``` If you have administrator privileges, you can make **swarm** accessible for all users. Simply copy the binary to `/usr/bin/`. The man page can be installed this way: accessible for all users. Simply copy the binary `./bin/swarm` to `/usr/local/bin/` or to `/usr/bin/`. The man page can be installed this way: ```sh cd ./man/ Loading Loading @@ -210,15 +223,10 @@ from two different sets have the same hash code, it means that the sequences they represent are identical. If for some reason your fasta entries don't have abundance values, and you still want to run swarm, you can easily add fake abundance values: ```sh sed '/^>/ s/$/_1/' amplicons.fasta > amplicons_with_abundances.fasta ``` Alternatively, you may specify a default abundance value with **swarm**'s `--append-abundance` (`-a`) option to be used when abundance information is missing from a sequence. you still want to run swarm (not recommended), you can specify a default abundance value with **swarm**'s `--append-abundance` (`-a`) option to be used when abundance information is missing from a sequence. ### Launch swarm ### Loading Loading @@ -305,15 +313,6 @@ rm "${AMPLICONS}" ``` ## Troubleshooting ## If **swarm** exits with an error message saying `This program requires a processor with SSE2`, your computer is too old to run **swarm** (or based on a non x86-64 architecture). **swarm** only runs on CPUs with the SSE2 instructions, i.e. most Intel and AMD CPUs released since 2004. ## Citation ## To cite **swarm**, please refer to: Loading @@ -333,7 +332,7 @@ You are welcome to: * submit suggestions and bug-reports at: https://github.com/torognes/swarm/issues * send a pull request on: https://github.com/torognes/swarm/ * compose a friendly e-mail to: Frédéric Mahé <mahe@rhrk.uni-kl.de> and Torbjørn Rognes <torognes@ifi.uio.no> * compose a friendly e-mail to: Frédéric Mahé <frederic.mahe@cirad.fr> and Torbjørn Rognes <torognes@ifi.uio.no> ## Third-party pipelines ## Loading @@ -356,7 +355,7 @@ You are welcome to: If you want to try alternative free and open-source clustering methods, here are some links: * [VSEARCH](https://github.com/torognes/vsearch) * [vsearch](https://github.com/torognes/vsearch) * [Oligotyping](http://merenlab.org/projects/oligotyping/) * [DNAclust](http://dnaclust.sourceforge.net/) * [Sumaclust](http://metabarcoding.org/sumatra) Loading @@ -365,6 +364,11 @@ methods, here are some links: ## Version history ## ### version 3.0 ### **swarm** 3.0 is much faster when _d_ = 1, and consumes less memory. Strict dereplication is now mandatory. ### version 2.2.2 ### **swarm** 2.2.2 fixes a bug causing Swarm to wait forever in very rare Loading man/swarm.1 +215 −86 File changed.Preview size limit exceeded, changes collapsed. Show changes scripts/amplicon_contingency_table.py +7 −9 Original line number Diff line number Diff line #!/usr/bin/env python #!/usr/bin/env python3 # -*- coding: utf-8 -*- """ Read all fasta files and build a sorted amplicon contingency table. Usage: python amplicon_contingency_table.py samples_*.fas table. Usage: python3 amplicon_contingency_table.py samples_*.fas """ from __future__ import print_function __author__ = "Frédéric Mahé <mahe@rhrk.uni-kl.fr>" __date__ = "2016/03/12" __version__ = "$Revision: 2.1" __author__ = "Frédéric Mahé <frederic.mahe@cirad.fr>" __date__ = "2019/09/24" __version__ = "$Revision: 3.0" import os import sys Loading @@ -35,7 +33,7 @@ def fasta_parse(): sample = os.path.basename(fasta_file) sample = os.path.splitext(sample)[0] samples[sample] = samples.get(sample, 0) + 1 with open(fasta_file, "rU") as fasta_file: with open(fasta_file, "r") as fasta_file: for line in fasta_file: if line.startswith(">"): amplicon, abundance = line.strip(">;\n").split(separator) Loading Loading @@ -65,7 +63,7 @@ def main(): all_amplicons, amplicons2samples, samples = fasta_parse() # Sort amplicons by decreasing abundance (and by amplicon name) sorted_all_amplicons = sorted(all_amplicons.iteritems(), sorted_all_amplicons = sorted(iter(all_amplicons.items()), key=operator.itemgetter(1, 0)) sorted_all_amplicons.reverse() Loading Loading
.travis.yml 0 → 100644 +15 −0 Original line number Diff line number Diff line language: c++ os: linux dist: bionic compiler: gcc before_install: - sudo apt-get install -y valgrind script: - make - export PATH=$PWD/bin:$PATH - git clone https://github.com/frederic-mahe/swarm-tests.git && cd swarm-tests && bash ./run_all_tests.sh | tee tests.log && ! grep -q FAIL tests.log
Makefile 0 → 100644 +37 −0 Original line number Diff line number Diff line # SWARM # # Copyright (C) 2012-2019 Torbjorn Rognes and Frederic Mahe # # This program is free software: you can redistribute it and/or modify # it under the terms of the GNU Affero General Public License as # published by the Free Software Foundation, either version 3 of the # License, or (at your option) any later version. # # This program is distributed in the hope that it will be useful, # but WITHOUT ANY WARRANTY; without even the implied warranty of # MERCHANTABILITY or FITNESS FOR A PARTICULAR PURPOSE. See the # GNU Affero General Public License for more details. # # You should have received a copy of the GNU Affero General Public License # along with this program. If not, see <http://www.gnu.org/licenses/>. # # Contact: Torbjorn Rognes <torognes@ifi.uio.no>, # Department of Informatics, University of Oslo, # PO Box 1080 Blindern, NO-0316 Oslo, Norway # Makefile for SWARM PROG=bin/swarm MAN=man/swarm.1 swarm : $(PROG) $(PROG) : make -C src swarm install : $(PROG) $(MAN) /usr/bin/install -c $(PROG) '/usr/local/bin' /usr/bin/install -c $(MAN) '/usr/local/share/man/man1' clean : make -C src clean
README.md +37 −33 Original line number Diff line number Diff line [](https://travis-ci.org/torognes/swarm) # swarm A robust and fast clustering method for amplicon-based studies. Loading @@ -16,7 +18,18 @@ To help users, we describe starting from raw fastq files, clustering with **swarm** and producing a filtered OTU table. swarm 2.0 introduces several novelties and improvements over swarm swarm 3.0 introduces: * a much faster default algorithm, * a reduced memory footprint, * binaries for Windows x86-64, GNU/Linux ARM 64, and GNU/Linux POWER8, * an updated, hardened, and thoroughly tested code. Please note that: * strict dereplication of input sequences is now mandatory, * \-\-seeds option (\-w) now outputs results sorted by decreasing abundance, and then by alphabetical order of sequence labels. swarm 2.0 introduced several novelties and improvements over swarm 1.0: * built-in breaking phase now performed automatically, * possibility to output OTU representatives in fasta format (option Loading @@ -24,13 +37,13 @@ swarm 2.0 introduces several novelties and improvements over swarm * fast algorithm now used by default for *d* = 1 (linear time complexity), * a new option called *fastidious* that refines *d* = 1 results and reduces the number of small OTUs, reduces the number of small OTUs. ## Common misconceptions **swarm** is a single-linkage clustering method, with some superficial similarities with other clustering methods (e.g., [Huse et al, 2010](http://www.ncbi.nlm.nih.gov/pmc/articles/PMC2909393/)). **swarm**'s similarities with other clustering methods (e.g., [Huse et al, 2010](http://www.ncbi.nlm.nih.gov/pmc/articles/PMC2909393/)). **swarm**'s novelty is its iterative growth process and the use of sequence abundance values to delineate OTUs. **swarm** properly delineates large OTUs (high recall), and can distinguish OTUs with as little as Loading Loading @@ -76,18 +89,18 @@ cgtcgtcgtcgtcgt where sequence identifiers are unique and end with a value indicating the number of occurrences of the sequence (e.g., `_1000`). Alternative format is possible with the option `-z`, please see the [user manual](https://github.com/torognes/swarm/blob/master/man/swarm_manual.pdf). Swarm format is possible with the option `-z`, please see the [user manual](https://github.com/torognes/swarm/blob/master/man/swarm_manual.pdf). Swarm **requires** each fasta entry to present a number of occurrences to work properly. That crucial information can be produced during the [dereplication](#dereplication) step. [dereplication](#dereplication-mandatory) step. Use `swarm -h` to get a short help, or see the [user manual](https://github.com/torognes/swarm/blob/master/man/swarm_manual.pdf) for a complete description of input/output formats and command line options. The memory footprint of **swarm** is roughly 1.6 times the size of the The memory footprint of **swarm** is roughly 0.6 times the size of the input fasta file. When using the fastidious option, memory footprint can increase significantly. See options `-c` and `-y` to control and cap swarm's memory consumption. Loading @@ -105,14 +118,14 @@ using the ```sh git clone https://github.com/torognes/swarm.git cd swarm/src/ cd swarm/ make cd ../bin/ ``` If you have administrator privileges, you can make **swarm** accessible for all users. Simply copy the binary to `/usr/bin/`. The man page can be installed this way: accessible for all users. Simply copy the binary `./bin/swarm` to `/usr/local/bin/` or to `/usr/bin/`. The man page can be installed this way: ```sh cd ./man/ Loading Loading @@ -210,15 +223,10 @@ from two different sets have the same hash code, it means that the sequences they represent are identical. If for some reason your fasta entries don't have abundance values, and you still want to run swarm, you can easily add fake abundance values: ```sh sed '/^>/ s/$/_1/' amplicons.fasta > amplicons_with_abundances.fasta ``` Alternatively, you may specify a default abundance value with **swarm**'s `--append-abundance` (`-a`) option to be used when abundance information is missing from a sequence. you still want to run swarm (not recommended), you can specify a default abundance value with **swarm**'s `--append-abundance` (`-a`) option to be used when abundance information is missing from a sequence. ### Launch swarm ### Loading Loading @@ -305,15 +313,6 @@ rm "${AMPLICONS}" ``` ## Troubleshooting ## If **swarm** exits with an error message saying `This program requires a processor with SSE2`, your computer is too old to run **swarm** (or based on a non x86-64 architecture). **swarm** only runs on CPUs with the SSE2 instructions, i.e. most Intel and AMD CPUs released since 2004. ## Citation ## To cite **swarm**, please refer to: Loading @@ -333,7 +332,7 @@ You are welcome to: * submit suggestions and bug-reports at: https://github.com/torognes/swarm/issues * send a pull request on: https://github.com/torognes/swarm/ * compose a friendly e-mail to: Frédéric Mahé <mahe@rhrk.uni-kl.de> and Torbjørn Rognes <torognes@ifi.uio.no> * compose a friendly e-mail to: Frédéric Mahé <frederic.mahe@cirad.fr> and Torbjørn Rognes <torognes@ifi.uio.no> ## Third-party pipelines ## Loading @@ -356,7 +355,7 @@ You are welcome to: If you want to try alternative free and open-source clustering methods, here are some links: * [VSEARCH](https://github.com/torognes/vsearch) * [vsearch](https://github.com/torognes/vsearch) * [Oligotyping](http://merenlab.org/projects/oligotyping/) * [DNAclust](http://dnaclust.sourceforge.net/) * [Sumaclust](http://metabarcoding.org/sumatra) Loading @@ -365,6 +364,11 @@ methods, here are some links: ## Version history ## ### version 3.0 ### **swarm** 3.0 is much faster when _d_ = 1, and consumes less memory. Strict dereplication is now mandatory. ### version 2.2.2 ### **swarm** 2.2.2 fixes a bug causing Swarm to wait forever in very rare Loading
scripts/amplicon_contingency_table.py +7 −9 Original line number Diff line number Diff line #!/usr/bin/env python #!/usr/bin/env python3 # -*- coding: utf-8 -*- """ Read all fasta files and build a sorted amplicon contingency table. Usage: python amplicon_contingency_table.py samples_*.fas table. Usage: python3 amplicon_contingency_table.py samples_*.fas """ from __future__ import print_function __author__ = "Frédéric Mahé <mahe@rhrk.uni-kl.fr>" __date__ = "2016/03/12" __version__ = "$Revision: 2.1" __author__ = "Frédéric Mahé <frederic.mahe@cirad.fr>" __date__ = "2019/09/24" __version__ = "$Revision: 3.0" import os import sys Loading @@ -35,7 +33,7 @@ def fasta_parse(): sample = os.path.basename(fasta_file) sample = os.path.splitext(sample)[0] samples[sample] = samples.get(sample, 0) + 1 with open(fasta_file, "rU") as fasta_file: with open(fasta_file, "r") as fasta_file: for line in fasta_file: if line.startswith(">"): amplicon, abundance = line.strip(">;\n").split(separator) Loading Loading @@ -65,7 +63,7 @@ def main(): all_amplicons, amplicons2samples, samples = fasta_parse() # Sort amplicons by decreasing abundance (and by amplicon name) sorted_all_amplicons = sorted(all_amplicons.iteritems(), sorted_all_amplicons = sorted(iter(all_amplicons.items()), key=operator.itemgetter(1, 0)) sorted_all_amplicons.reverse() Loading