Access request submitted ✓
Private beta — now accepting researchers

AI native platform for multi-omics research

OmniBioAI unifies reproducible multi-omics analysis, agentic AI reasoning, and enterprise-grade infrastructure — across local, HPC, and cloud environments.

↗ Request Access ▶ Watch Demo 🌐 Try Web App
QR code for OmniBioAI Web Studio Open on your phone Scan to launch Web Studio
12,000+TES tool definitions
869Workflows
500+Workbench plugins
1,320+Container baseline
40+Agentic Pipelines/Plugins
28Microservices
28M+Abstracts
150+Domains
12Organisms
Live platform metrics & health →

// Biomedical AI — Literature Intelligence

Powered by PubMed.
Enhanced by AI.

Grounded in evidence. Retrieval-augmented generation over 28.1M PubMed abstracts, running on your own hardware with Llama 3, and every answer cites the PubMed IDs it came from.

01
Question A natural-language biomedical question, optionally scoped to one research domain
02
Query expansion (LLM) The question is rewritten into related variants so relevant papers are not missed on wording alone
03
Vector search (FAISS) Each variant is embedded with mxbai-embed-large (1024-d) and matched against per-domain FAISS indexes; hits are merged by PMID
04
Hybrid keyword search (BM25) Within a named domain, BM25 keyword results are fused with the vector results
05
Re-ranking Candidate abstracts are re-scored so the most relevant evidence reaches the model
06
Knowledge graph enrichment (Neo4j) Optional graph lookup adds linked biomedical entities as context
07
Answer generation (Llama 3) Llama 3 runs locally through Ollama and writes the answer from the retrieved abstracts only
08
PMID citations Every passage carries its PubMed ID, so each claim can be traced back to the paper
28.1M Distinct PubMed abstracts stored locally
282 Indexed domains: 150 research topics + 132 general-corpus shards
75.3M Vectors across the FAISS indexes, being re-embedded to 1024-d with mxbai-embed-large
Local Llama 3 via Ollama: questions and data stay on your server
Llama 3 via Ollama mxbai-embed-large FAISS + BM25 hybrid Neo4j knowledge graph LLM query expansion Re-ranking PMID-cited answers Per-domain indexes Runs locally · private

Indexes are being re-embedded to 1024 dimensions; a domain becomes queryable as its index is rebuilt. Live re-indexing progress ↗ · How the pipeline works ↗


// Core capabilities

Built for real science,
not demos

Every component is designed with reliability and traceability at its core.

⬡
Agentic AI reasoning

LangGraph-based orchestration with RAG-powered assistants using Hugging Face and Ollama — bridging deterministic bio-computation with explainable AI interpretation.

◈
Workflow agnostic

Native support for WDL, Nextflow, Snakemake, and CWL. Cloud-agnostic Tool Execution Service ensures 1:1 parity between local and massive-scale genomic execution.

◎
GPU-accelerated stack

Optimized on NVIDIA PyTorch with CUDA on DGX Spark. High-performance GNNs for drug discovery and deep learning for single-cell transcriptomics.

⬕
End-to-end provenance

Production-grade Model Registry and LIMS-X metadata system for traceability — from raw FASTQ to drug-target intelligence and pathway enrichment.

Lab Integrations:
· Benchling (samples, sequences, entities, notebooks)
· REDCap (patient cohorts, clinical research data)
· eLabFTW (experiment export, file upload, tags)
· AWS S3 (import datasets, export results, MinIO compatible)
· Zenodo (publish with DOI, open science, reproducible research)
· LIMS built-in
· ELN-ready export: PDF reports (WeasyPrint) · ELN JSON (machine-readable) · Benchling notebook export · eLabFTW experiment export · CSV/TSV universal format

🔐
SSO Support

Sign in with Google, GitHub, or Microsoft, alongside secure email/password authentication via JWT + zero-trust auth service. Existing accounts can link a provider identity with a one-time confirmation step.

◇
Reliability-first design

Automated test suites across the microservices, with short-lived tokens, role-based access control and an append-only audit log built in from day one.

⬡
Plugin ecosystem

Extensible plugin architecture with 500+ Workbench plugins and an OnboardAI documentation browser. Documentation and plugin inventories vary by repository release.

⚙
Modular pipeline system

Each bioinformatics workflow is modular, allowing researchers to customize, extend, or replace analysis steps without breaking reproducibility.

☁
Cloud & HPC ready

Runs consistently across local machines, Slurm-based HPC clusters, and cloud platforms like AWS, Azure, and GCP with identical execution logic.

🐳
Docker Images

*Historical beta-release inventory; live availability varies by deployment.
320 Docker images + 1,000 ARM64 SIF images in that baseline
Available: ghcr.io/omnibioai & huggingface.co/omnibioai
ARM64 + x86_64 support

🤖
Agentic workflow orchestration

LangGraph-based agentic workflows dynamically plan, execute, and adapt bioinformatics pipelines, enabling multi-step reasoning across tools, datasets, and analysis stages.

🛡
Production-Grade Observability

Full-stack error tracking and performance monitoring via Sentry.io across all 28 microservices — with release tracking, error alerting, and performance tracing built in from day one.

⟐
Real-Time Collaboration

WebSocket-based shared workspaces with live chat, object-level comments, annotation history, and an audit log of security-relevant events.

🎓
Zenodo Integration

Publish analysis results to Zenodo with automatic DOI generation for citable, reproducible research. Search and import public datasets directly from Zenodo.

📓
eLabFTW Integration

Export analysis results directly to eLabFTW electronic lab notebook. Creates experiment entries, uploads output files, and adds tags automatically. Open-source ELN used by thousands of academic labs worldwide.


// Enterprise & scientific workflow integrations

Scientific work,
connected across your stack

OmniBioAI's agentic workflows connect research, data, compute, MLOps, cloud infrastructure, and enterprise systems — so scientists and AI agents can work across the tools your organization already uses, without leaving the platform.

67 integrations across 9 categories Browse the full catalog in the docs ↗

Genomics & Bio Platforms

9

Sequencing, cloud-genomics and bioinformatics platforms, plus GA4GH interoperability.

Workflow, Orchestration & Compute

10

Workflow engines, pipeline orchestrators, cluster schedulers and notebook environments.

Cloud Storage & Data Transfer

11

Object storage, file sharing, private cloud and bulk research-data transfer.

Data Warehouses & Analytics

4

Warehouse and lakehouse platforms and the tooling that models data inside them.

ML & Experiment Tracking

4

Experiment tracking, model hubs and data/model versioning.

LIMS, ELN & Biobank

7

Laboratory information management, electronic lab notebooks and biospecimen management.

Data Repositories & Metadata

4

Research data repositories and data-catalog / metadata platforms.

CI/CD, Source Control & Containers

10

Git hosting, build and pipeline systems, and container registries.

Collaboration, ITSM & Operations

8

Messaging, Microsoft 365, issue and incident management, monitoring and secrets health.

// Biomedical reference databases

Reference knowledge,
one query away

Reference-database plugins give scientists and AI agents read-only access to the genomic, clinical, protein, chemical and ontology resources that analyses depend on, from variant interpretation to target discovery.

130 reference databases across 9 categories · 100 implemented · 30 scaffolded Browse the full catalog in the docs ↗

Genes, Variants & Clinical Genetics

40

Gene nomenclature and annotation, population and structural variation, and clinical-genetics evidence.

1000 Genomes / IGSRPopulation samples, data collections and files CIViCCancer variant evidence items and assertions ClinGenGene–disease validity and dosage sensitivity ClinVarClinical significance of human variants COSMICCancer Gene Census via licensed downloads dbNSFPFunctional predictions for missense variants dbSNPrsIDs, alleles and genomic placements dbVarNCBI structural-variation studies and variants DECIPHERInterface metadata and connectivity only DGVStructural variants from validated local snapshots EnsemblGenes, transcripts, sequences, variants and homology Gene2PhenotypeCurated gene–disease panels from EMBL-EBI Genomics England PanelAppExpert-reviewed diagnostic gene panels gnomADPopulation allele frequencies and gene constraint GWAS CatalogStudies, traits and reported associations HGMDLicensed clinical variants via a local or gateway provider HGNCApproved human gene symbols and names MARRVELRare-variant evidence across model organisms MedGenMedical-genetics concepts, genes and features NCBI EntrezGene, protein, nucleotide and taxonomy search OMIMMendelian genes, disorders and inheritance PharmGKBPharmacogenomic annotations via ClinPGx PharmVarPharmacogene star alleles and haplotypes PheGenIPhenotype–genotype associations from NCBI SNPediaCommunity variant and genotype annotations UCSC Genome BrowserGenome databases, tracks and region data Clinvitae scaffoldAggregated clinical variant reports CPIC scaffoldPharmacogenomic prescribing guidelines GeneCards scaffoldIntegrated human gene summaries GeneReviews scaffoldExpert-authored disease reviews GWAS Summary Statistics scaffoldFull summary statistics for GWAS LOVD scaffoldLocus-specific variant databases Mastermind scaffoldVariant evidence from the literature MaveDB scaffoldMultiplexed assays of variant effect NCBI Gene scaffoldGene records and annotation OncoKB scaffoldPrecision-oncology variant knowledge RefSeq scaffoldNCBI reference sequences TOPMed scaffoldDeep-sequencing variant frequencies VarCards scaffoldIntegrated coding-variant annotation VarSome scaffoldVariant classification and annotation

Expression, Omics & Cohorts

19

Sequencing archives, expression atlases, samples, cell lines and cancer cohorts.

Proteins, Structure & Interactions

18

Protein knowledgebases, families, 3D structures, orthology, proteomics and interaction networks.

Drugs, Chemistry & Metabolomics

20

Drugs and targets, chemical compounds, enzymes, lipids, metabolomics and toxicogenomics.

Ontologies, Disease & Phenotype

12

Biomedical ontologies, phenotype vocabularies and gene–disease association resources.

Model Organisms

7

Model-organism databases for fly, mouse, rat, yeast, plant, worm and zebrafish.

Non-coding RNA & Regulation

6

Regulatory elements, transcription-factor binding profiles and non-coding RNA.

Pathways

6

Curated pathway and metabolic-network knowledgebases.

Clinical Trials & Models

2

Resources that span domains: clinical-trial registries and biomolecular AI models.

Entries tagged scaffold are registered but not yet connected to their upstream resource; the other 100 call it in source. A few connectors report availability status only, where the upstream offers no usable API. Licensed resources (for example COSMIC, DrugBank, HGMD) need your own licence and credentials. Counts are from the docs catalog generated 2026-09-19, updated for connectors implemented since; no entry is asserted as deployed or operationally verified.


// Catalogs

What's in the platform

All catalogs in the docs ↗
869WorkflowsIn 102 domainsEngine mix in docs catalog: Nextflow 861 · WDL 5 · Snakemake 2 · CWL 1 501Workbench pluginsAcross 22 categories, including 66 integrations 12,276TES tool definitions8 execution backends · 9,061 auto-imported and not yet verified 130Reference database plugins100 implemented · 30 scaffolded 219API routesFound in source across 8 services · 109 documented publicly

Plugin, tool, reference database and API counts are read from the documentation's generated catalogs (19 Sep 2026). They describe what is registered or configured in source, not that every entry has been tested or deployed. How to read these figures ↗


// Multi-omics coverage

Every modality,
one platform

From raw sequencing reads to biological insight — 869 workflows across 102 research domains, grouped here into the areas they serve.

Genomics
Genome & variant analysis

Germline variant calling, annotation and prioritisation, from exomes to assemblies.

WGSWESSV callingML variant prioritisationGenome assemblyPangenomeOptical mappingMitochondrialTelomere
Long-read
Long-read sequencing

Nanopore and PacBio alignment, assembly, methylation and isoform analysis.

NanoporePacBioNative methylationIsoforms
Transcriptomics
Bulk & specialised RNA

Expression, splicing and RNA biology beyond standard RNA-seq.

RNA-seqLong-read RNAmiRNAlncRNAcircRNARibo-seqRNA editing
Single-cell
Single-cell & multiome

QC, integration, clustering, trajectories and multimodal single-cell assays.

scRNA-seqCell RangerCITE-seqscATAC-seqMultiome / ArchR
Spatial
Spatial omics

Spatially resolved expression and protein maps with cell-type deconvolution.

Visium / HDXeniumCODEX / PhenoCyclerSpatial proteomicsSpatial multi-omics
Epigenomics
Functional genomics & epigenomics

Regulatory elements, chromatin state, 3D genome and CRISPR screens.

ATAC-seqChIP-seqMethylationHi-CCRISPR screensBase editing
Proteomics · Metabolomics
Proteomics & metabolomics

Mass-spec quantification, enrichment and metabolite profiling.

Mass specProteogenomicsMetabolomicsExposomeGSEA / GSVA
Multi-omics
Multi-omics integration

Joint analysis across genomic, transcriptomic, epigenomic and proteomic layers.

Multimodal integrationMulti-omics QCCross-layer methods
Oncology
Cancer genomics

Somatic mutations, tumour burden and liquid biopsy analysis.

Somatic callingTMB / MSIctDNAFragmentomics
Immunology
Immunology & infectious disease

Immune profiling, HLA and repertoire analysis, pathogens and vaccines.

HLA typingVDJ repertoireImmune deconvolutionVaccine designViral genomics
Microbiome
Microbiome & metagenomics

Taxonomic and functional profiling, including host–microbe interactions.

Shotgun metagenomicsTaxonomic profilingHost–microbiome
Clinical
Clinical & translational

Cohort analysis, clinical reporting and precision-medicine workflows, with REDCap import and export.

Clinical reportingPharmacogenomicsPolygenic riskClinical NLPTrial matchingTransplantReproductive
Drug discovery
Drug discovery

GNN-based target identification, drug response and repurposing.

Target identificationDrug repurposingDrug synergyResponse predictionADMET
Population
Population & specialised genomics

Population, evolutionary and ancient genomes, plus neuro and aging studies.

Population geneticsEvolutionaryAncient DNAForensicsNeurogenomicsAgingEmerging methods
AI · Platform
AI models & platform methods

Foundation models, knowledge graphs and the tooling that keeps pipelines reproducible.

scGPT / foundation modelsBiomedical knowledge graphBenchmarkingSimulationPipeline chainingReference data

Browse every workflow domain in the docs ↗

PythonRNextflow WDLSnakemakeCWL GATKSeuratDESeq2 LangGraphPyTorchCUDA DockerKubernetesFastAPI DjangoCeleryRedis MySQLAWSAzureGCP SlurmHuggingFaceOllamaSentry BenchlingREDCapeLabFTW AWS S3Zenodo

// Reference genomes

Reference Genomes

12 species with genome sequence and gene annotation on disk, and transcript sequences for 11 of them. Prepared aligner indexes are ready for human (STAR, Cell Ranger) and mouse (Bowtie2, Cell Ranger).

Species Common Name Assemblies Annotation Use Cases On disk
Homo sapiens Human GRCh38 (hg38), hg19 GENCODE v44, v46 WGS, WES, scRNA, ATAC, clinical GenomeTranscriptsGTFSTARCell Ranger
Mus musculus Mouse GRCm39, mm10 GENCODE vM33 Drug discovery, knockout models GenomeTranscriptsGTFBowtie2Cell Ranger
Rattus norvegicus Rat GRCr8 Ensembl 111 Pharmacology, toxicology GenomeGTF
Pan troglodytes Chimpanzee Pan_tro_3.0 Ensembl 112 Comparative genomics, evolution, primate research GenomeTranscriptsGTF
Macaca mulatta Rhesus Macaque Mmul_10 Ensembl 112 Primate translational research GenomeTranscriptsGTF
Danio rerio Zebrafish GRCz11 Ensembl 112 Developmental biology, screens GenomeTranscriptsGTF
Drosophila melanogaster Fruit Fly BDGP6 Ensembl 112 Genetics, CRISPR screens GenomeTranscriptsGTF
Saccharomyces cerevisiae Yeast R64-1-1 Ensembl 112 Yeast two-hybrid, metabolomics GenomeTranscriptsGTF
Caenorhabditis elegans C. elegans WBcel235 Ensembl 112 Aging, longevity, RNAi screens, neuroscience GenomeTranscriptsGTF
Arabidopsis thaliana Thale Cress TAIR10 Ensembl Plants 59 Plant biology, epigenomics, stress response GenomeTranscriptsGTF
Sus scrofa Pig Sscrofa11.1 Ensembl 112 Xenotransplantation, cardiovascular, metabolic disease GenomeTranscriptsGTF
Gallus gallus Chicken GRCg7b Ensembl 112 Immunology, developmental biology, vaccine research GenomeTranscriptsGTF
STAR Prepared index, ready to use GTF Sequence or annotation file

Inventory of local reference storage, measured 19 Sep 2026. Live status on the platform dashboard ↗ · Full inventory in the docs ↗


// System requirements

What you need
to get started

OmniBioAI Studio runs on any modern Linux or macOS machine with Docker installed. One download, no command line needed.

💾
Memory

Minimum: 16 GB RAM
Recommended: 32 GB RAM
With local LLM: 64 GB RAM

💿
Storage

Installer: ~75–105 MB
Docker images: ~10 GB (one-time pull)
Data + work dirs: 50–200 GB

🖥️
Operating System

Linux: Ubuntu 20.04+ (AppImage)
macOS: 12+ Apple Silicon + Intel
Windows: WSL2 + Docker

🐳
Docker

Docker Engine 24+ or Docker Desktop
Docker Compose v2 (included)
No other dependencies required

⚡
GPU (Optional)

NVIDIA GPU + nvidia-container-toolkit
Required only for local LLM inference
Cloud API (Claude/GPT) works without

🌐
Network

Internet required for first boot only
~10 GB pulled from ghcr.io automatically
Fully offline after first run

🔑
License

30-day free trial included
License key with every download
7-day offline grace period
[email protected]

🧬
jq

Required: sudo apt install jq
Or: brew install jq
Used for health checks and config
Included in most Linux distros

● Private Beta · v0.7.1-beta

Access requires approval

Approved researchers receive a platform-specific download link and onboarding support within 1–2 business days.

↗ Request Beta Access

License key included · 30-day free trial


// Downloads

OmniBioAI Studio v0.7.1-beta

Choose your platform and architecture. All installers include a 30-day free trial license.

🐧Linux x86_64Intel / AMD · Ubuntu 20.04+
🦾Linux ARM64DGX · Graviton · Raspberry Pi
🪟WindowsWindows 10 / 11 x64 · WSL2
Linux install commands
sudo dpkg -i omnibioai-studio_*amd64.deb sudo rpm -i omnibioai-studio-*.x86_64.rpm sudo dpkg -i omnibioai-studio_*arm64.deb sudo rpm -i omnibioai-studio-*.aarch64.rpm chmod +x *.AppImage && ./OmniBioAI*.AppImage
All releases on GitHub ↗

// System architecture

Built for sensitive data

A containerized, microservices-led environment built for massive scale, traceability, and high performance.

Engineered for provenance

OmniBioAI Studio separates user experience orchestration from heavy-lifting workflow engines. It executes multi-omics routines natively, passing biological insights directly into automated reporting and visualization pipelines.

With isolation between its core service layers, researchers can deploy pipelines locally or easily burst to Slurm or cloud systems without code modifications.

↗ See live architecture & health
8 Execution backends
219 API routes · 8 services
12 Organisms with references
282 Literature indexes
Execution backends
LocalSlurm (DGX)KubernetesAWS BatchAzure BatchGCP BatchHTTP ToolServer
Backend definitions in the docs ↗
OmniBioAI Studio UI (Desktop Frontend)
│ Authenticated API calls
Agentic AI Orchestration (LangGraph / Ollama / HF)
│ Orchestration & data mapping
BioFlow Runtime Engine (Nextflow / WDL / Snakemake)
│ Tracking & provenance link
LIMS-X Metadata & Sample Tracking System
│ Infrastructure layer
GPU Accelerated Stack (CUDA / NVIDIA DGX Spark)
HIPAA-aligned · PHI protection

PHI stays on your infrastructure

OmniBioAI Studio is built with HIPAA-aligned security controls for protected health information (PHI) and secure data handling. It runs locally on your own machines, so sequencing data, clinical metadata and results stay in your environment by default, and anything sent to an external service is explicitly enabled and reviewed. It is designed to support hospitals, clinical labs and research institutions in their own HIPAA compliance programs.

Local-first data PHI is stored and processed on your hardware by default, not ours.
On-device AI Local models via Ollama and Hugging Face keep prompts and data in-house.
Access control Short-lived tokens, organization SSO, OAuth and MFA, plus role-based access control with org-scoped data.
Tamper-resistant audit log Logins, permission denials, role changes and policy decisions go to an append-only log that application credentials cannot alter, with legal-hold support.

Every security, testing and reproducibility claim in our documentation carries an evidence label (configured, tested, deployed or operationally verified), and known limits are published rather than hidden. Quality & Trust ↗ · Security model ↗


// Academic presence

Peer-reviewed research

Methods developed and powered by OmniBioAI platform architectures across transcriptomics and proteomics.

2025
Genetic mutations in lymphocytic variant of hypereosinophilic syndrome: study of five siblings
Frontiers in Medicine · December 2025
↗ View paper
2018
Whole Exome Sequencing identifies common and rare variant Metabolic QTLs in a Middle Eastern Population
Nature Communications · January 2018
↗ View paper
2015
MetaRNA-Seq: An Interactive Tool to Browse and Annotate Metadata from RNA-Seq Studies
BioMed Research International · August 2015
↗ View paper

// Apply for private beta

Request Access

Get access to OmniBioAI Studio v0.7.1-beta. Accelerate your multi-omics data integration with robust, explainable AI workflows.

Your Operating System *