SeqHub's Diversity Search widens your results to surface distant homologs that other search algorithms rank low or miss entirely. When you find a hit with low sequence identity, you can run a structural alignment within SeqHub to validate it before committing time to the candidate.
Tatta Bio
@tattabio.bsky.social
Building genomic intelligence Metagenomic datasets, genomic language models, SeqHub Our research: https://tatta.bio Analyze your sequences on https://seqhub.org
SeqHub has an updated interface and new features! Join a live walkthrough on August 11 at 11am ET to see SeqHub in action, whether you're just getting started or want to catch up on what's changed since our last update. Sign up: docs.google.com/forms/d/e/1F...
Retrieving gene neighborhoods is still one of the harder parts of working with sequence data, often needing large downloads and manual parsing of genomic locations. The SeqHub API returns functional annotations and surrounding genes for a query protein in one call, with locations already resolved.
SeqHub, now in dark mode. Curious about the new UI or want to get up to speed on the latest SeqHub features? Join us for a webinar on August 11 at 11am EST. Sign up link below.
We'll be at the AI x Bio Summit today at the NYSE, hosted by Decoding Bio. If you're at the event and curious to learn more about our models or platform, SeqHub, send us a DM or find Steph Flamen on the floor.
We just launched a new UI for SeqHub so things might look a little different next time you log in 👀 With your feedback, we rebuilt navigation across the platform, with most tools now accessible from one location.
We've arrived at ICML! @microyunha.bsky.social will be speaking at the GenBio Workshop on July 10th at 1:30pm local time: "Genomic Language Modeling for Context-Aware Biological Discovery." Hope to see you there! #ICML2026
This week we are at ICML in Seoul, and next week we're headed to Washington, D.C. for ISMB. Our ICML (GenBio Workshop) talk: July 10th at 1:30pm Our ISMB talk: July 15 at 12:30pm We'd love to connect. Send a DM or email team@tatta.bio and we'll find time!
In SeqHub, you can now search a protein to find FlashPPI2-predicted interaction partners across all 400 million proteins and 132K microbial genomes in our database (OpenGenome). You can still upload full genomes to find interactions within and between genomes.
Two weeks ago, our FlashPPI paper was published in @pnas.org. Today, we introduce our updated model, FlashPPI2. Fine-tuned on new AlphaFold structures, this new model achieves a 17% improvement over FlashPPI on the E. coli protein interaction benchmark. seqhub.org/blog/flashppi2
FlashPPI2: Enhanced Model, Scaled Across 130,000+ Genomes - SeqHub
FlashPPI2 drives up PPI prediction performance (AUPRC) by 17% while maintaining inference speed at minutes per genome — now deployed across SeqHub's database of over 130,000 microbial genomes.
seqhub.org
We'll be at ICML in Seoul July 6 - 11! Our Chief Scientist, @microyunha.bsky.social, will be speaking at the GenBio Workshop on July 10 at 1:30pm local time, presenting "Genomic Language Modeling for Context-Aware Biological Discovery." #ICML2026
"SeqHub has become an integral part of our workflow...[it's] typically the first place we go to begin understanding what a gene might be doing and to identify its genomic neighbors across bacterial genomes." - Jeremy Rock, Rockefeller University seqhub.org/blog/rock-la...
How the Rock Lab Uses SeqHub to Accelerate Discovery in Mtb and Mabs - SeqHub
The Rock Lab at Rockefeller University studies Mycobacterium tuberculosis and M. abscessus — two pathogens with large stretches of unannotated genome. SeqHub has become an integral part of their workf...
seqhub.org
Our FlashPPI paper is out in @pnas.org! FlashPPI is a model for proteome-wide protein-protein interaction prediction. 🔹4x better predictive performance & 2,400x faster than existing sequence-based methods 🔹20,000x faster than leading structure-based approaches. www.pnas.org/doi/10.1073/...
PNAS
Proceedings of the National Academy of Sciences (PNAS), a peer reviewed journal of the National Academy of Sciences (NAS) - an authoritative source of high-impact, original research that broadly spans...
pnas.org
Search a protein in SeqHub and relevant literature appears alongside your genomic context results. Using the PaperBLAST database (Price & Arkin 2024), SeqHub retrieves papers, with specific passages highlighted, citing the sequence you searched and closely related ones.
Last semester in MB 360: Scientific Inquiry in Microbiology at the Bench, a course designed by Dr. Carlos Goller and Camila Loyola, student teams extracted DNA, sequenced with Nanopore and Illumina, assembled genomes, and used SeqHub to annotate what they had assembled. seqhub.org/blog/nc-stat...
How an NC State Microbiology Course Used SeqHub to Turn Student Sequencing Data into Published Genomes - SeqHub
Dr. Carlos Goller and TA Camila Loyola built NC State's MB 360 around hands-on genome sequencing and annotation. Students extracted DNA, assembled Delftia genomes, and used SeqHub to annotate and publ...
seqhub.org
ASM Microbe starts today. We're at booth 2324, next to the Applied and Environmental Microbiology Hub. Come find Andre, Yunha, and Steph if you're curious to learn more, provide feedback, or just say hello!
Our Chief Scientist @microyunha.bsky.social is speaking at FOG tomorrow! Catch her talk on Genomic Language Modeling for Sequence Analysis and Management in the Age of AI at 12:30pm at the Biopharma Data Management Stage.
The SeqHub API (beta) is free for non-commercial use. You can start querying today, and for throughput beyond the current beta limits, reach out at team@tatta.bio. We look forward to hearing your feedback! Docs: docs.seqhub.org/introduction
We'd love to connect at ASM Microbe in D.C.! Stop by booth 2324 to learn more about our research or talk through how SeqHub could support your work.
With the SeqHub API (now in beta), you can infer function for hypothetical proteins, identify gene clusters, and reconstruct putative pathways in uncharacterized organisms. Each query pulls from 130,000+ microbial genomes. Request a token from your Dashboard. Docs: docs.seqhub.org/introduction
SeqHub API - Zudoku
The SeqHub API gives you programmatic access to protein search and annotation tools built on top of Tatta Bio's genomic language models.
docs.seqhub.org
Programmatic access to SeqHub has been one of the more common user requests. Today we're launching the SeqHub API in beta. Functional annotations for your protein and its genomic neighbors, derived from genomic language model embeddings. Free for non-commercial use. seqhub.org/blog/seqhub-...
The SeqHub API Is Now in Beta - SeqHub
The SeqHub API is now in public beta. Submit protein sequences programmatically and get back gLM2-powered functional annotations and genomic context — free for non-commercial use.
seqhub.org
Taxonomy-specific search is now live in SeqHub. By popular request: filter by species, genus, family, or order before running your search, and get more results for the organisms you care about. Works with CoSearch and Diversity Search too.
Collaboration in SeqHub: PIs reviewing annotations in-platform instead of through file handoffs, grad students and postdocs building shared datasets and crediting co-authors when they publish, and a class contributing to shared data under a faculty co-author. seqhub.org/blog/working...
Working Together in SeqHub - SeqHub
Collaboration is now live on SeqHub. Share datasets with labmates, assign read or write access, and add co-authors so everyone who contributed gets credited.
seqhub.org
Mechanistic insights in host-virus biology often start with identifying which proteins are interacting across that boundary. SeqHub supports cross-proteome protein-protein interaction prediction. Upload any two genome FASTAs and get back a predicted interaction network spanning both.
SeqHub can surface protein hits that wouldn't show up in a standard search, through embedding-based retrieval, CoSearch (co-occurrence search), and diversity control. Each targets a different gap in what alignment-based tools return. seqhub.org/blog/why-seq...
Why SeqHub Search Finds Biology Other Tools Miss - SeqHub
SeqHub uses embedding-based search, diversity control, and co-occurrence search to surface protein connections that BLAST and profile-based tools miss — especially for hypothetical proteins and underr...
seqhub.org
Non-coding RNA is now annotated automatically in the SeqHub genome viewer. Upload a genome and Rfam annotations surface directly on the gene map alongside your coding sequences. Rfam annotations are now available across search and genome upload.
SeqHub now generates a UMAP from your uploaded protein family, so you can find clusters, better navigate your dataset, and prioritize candidates for validation. Just select a protein in your data table to launch it.
We're offering free 1:1 SeqHub training sessions, tailored to your research. New to the platform or already using it, book a time and we'll walk through what's most relevant to your work. calendar.app.google/KLFsKvszMMko...
Add collaborators to share data you've curated in SeqHub! Invited collaborators can... - View functional annotations - Add their own annotations and experimental results - Work with the SeqHub Agent - Run FlashPPI against the data No longer need to download and send files back and forth.
We'll be at the MIT Microbiome Symposium this Friday, April 17! If you're interested in learning more about our research at Tatta Bio or want to see how SeqHub can support your work, come find us at our table or set up time to chat (DM or email us at team@tatta.bio).