Open Science Index

Find open-source science resources

Cross-domain directory aggregating tools, AI models, datasets, and research resources from bio.tools, Bioconductor, HuggingFace, curated GitHub awesome-lists, and more.

Filters

Domain

text-generation29
fill-mask17
text-classification9
feature-extraction8
image-text-to-text5
question-answering5
image-feature-extraction3
sentence-similarity3
tabular-classification3
image-segmentation2
other2
text-ranking2
(None)72

Language

Python76
(None)92

License

(None)168

Source

huggingface168

Type(1)

Software tool3076
Database2418
AI model168

Filters

Domain

text-generation29
fill-mask17
text-classification9
feature-extraction8
image-text-to-text5
question-answering5
image-feature-extraction3
sentence-similarity3
tabular-classification3
image-segmentation2
other2
text-ranking2
(None)72

Language

Python76
(None)92

License

(None)168

Source

huggingface168

Type(1)

Software tool3076
Database2418
AI model168

168 of 5,662 resources

Showing 1–50

JThomas-CoE/coe-gemma4-biology-mmlu_pro-14b-a4b-q4

by JThomas-CoE

text-generation

Base model: google/gemma-4-26b-it Architecture: MoE — 26B total / ≈4B active parameters (1 shared expert + 8 routed from a pool of 128 per MoE layer, 30 MoE layers) Method: Activation-directed expert surgery — 128 → 64 experts per layer (50% reduction) Quantization: Q4KM (≈9.7 GB on disk) Tags:…

↓581 week ago

nvidia/geneformer_V2_104M_CLcancer

by nvidia

## Description: Geneformer is a foundational transformer model pretrained on a large-scale corpus of single-cell transcriptomes to enable context-specific predictions in settings with limited data in network biology. This model version was continually pretrained on ~14 million cancer transcriptomes…

↓135 months ago

Verdugie/STEM-Oracle-27B

by Verdugie

text-generation

# or·a·cle /ˈôrəkəl/ — a source of wise counsel; one who provides authoritative knowledge. From Latin ōrāculum, meaning divine announcement. In computer science, an oracle is a black box that always returns the correct answer — you don't ask it how it knows, you ask and it answers.

↓1362 months ago

nvidia/geneformer_V2_316M

by nvidia

## Description: Geneformer is a foundational transformer model pretrained on a large-scale corpus of single-cell transcriptomes to enable context-specific predictions in settings with limited data in network biology.

↓265 months ago

nvidia/geneformer_V1_10M

by nvidia

## Description: Geneformer is a foundational transformer model pretrained on a large-scale corpus of single-cell transcriptomes to enable context-specific predictions in settings with limited data in network biology.

↓165 months ago

keejkrej/cellpose-cpsam-onnx

by keejkrej

image-segmentation

ONNX export of the Cellpose cpsam (Cellpose-SAM) model for cell segmentation in microscopy images.

↓03 months ago

darkknight25/deepseek-16b-medical-GPT

by darkknight25

text-generation

darkknight25/deepseek-16b-medical-GPT is a fine-tuned version of deepseek-ai/deepseek-l6b-moe-chat, optimized for medical question answering, reasoning, and clinical summarization using QLoRA and open-access healthcare datasets.

↓010 months ago

nvidia/AMPLIFY_350M

by nvidia

> [!NOTE] > This model has been optimized using NVIDIA's TransformerEngine > library. Slight numerical differences may be observed between the original model and the optimized > model. For instructions on how to install TransformerEngine, please refer to the > official documentation.

↓278 months ago

mradermacher/Dans-PersonalityEngine-V1.3.0-24b-i1-GGUF

by mradermacher

For a convenient overview and download list, visit our model page for this model.

↓38910 months ago

mradermacher/Dans-PersonalityEngine-V1.2.0-24b-i1-GGUF

by mradermacher

If you are unsure how to use GGUF files, refer to one of TheBloke's READMEs for more details, including on how to concatenate multi-part files.

↓6241 year ago

prov-gigatime/GigaTIME

by prov-gigatime

↓2795 months ago

SaltySander/MOSAIC

by SaltySander

↓011 months ago

google/medasr

by google

automatic-speech-recognition

↓12.2K3 weeks ago

cambridgeltl/SapBERT-from-PubMedBERT-fulltext

by cambridgeltl

feature-extraction

datasets: - UMLS

↓1.8M2 years ago

nvidia/geneformer_V2_104M

by nvidia

## Description: Geneformer is a foundational transformer model pretrained on a large-scale corpus of single-cell transcriptomes to enable context-specific predictions in settings with limited data in network biology.

↓245 months ago

openadmet/pxr-chemeleon-baseline

by openadmet

> [!WARNING] > This is a baseline model trained on publicly available data. While we've done our best to curate the data, the model performance is quite poor. Proceed with caution.

↓262 weeks ago

nvidia/AMPLIFY_120M

by nvidia

> [!NOTE] > This model has been optimized using NVIDIA's TransformerEngine > library. Slight numerical differences may be observed between the original model and the optimized > model. For instructions on how to install TransformerEngine, please refer to the > official documentation.

↓6808 months ago

Keylab/COMO

by Keylab

COMO (Closed-loop Optical Molecule recOgnition) is a deep learning framework for Optical Chemical Structure Recognition (OCSR). It recognizes chemical structure diagrams from images and predicts SMILES strings with atom-level 2D coordinates and bond matrices.

↓01 day ago

ibm-research/biomed.omics.bl.sm.ma-ted-458m.moleculenet_bbbp

by ibm-research

Drugs targeting the central nervous system must meet stringent criteria for both efficacy and safety, including their ability to penetrate the blood-brain barrier (BBB). This model predicts the likelihood of small-molecule drugs crossing the BBB, a critical factor in CNS drug development.

↓341 year ago

ibm-research/biomed.omics.bl.sm.ma-ted-458m.tcr_epitope_bind

by ibm-research

T-cell receptor (TCR) binding to immunogenic peptides (epitopes) presented by major histocompatibility complex (MHC) molecules is a critical mechanism in the adaptive immune system, essential for antigen recognition and triggering immune responses.

↓451 year ago

ibm-research/biomed.omics.bl.sm.ma-ted-458m.protein_solubility

by ibm-research

Protein solubility is a critical factor in both pharmaceutical research and production processes, as it can significantly impact the quality and function of a protein. This is an example for finetuning ibm/biomed.omics.bl.sm-ted-458m for protein solubility prediction (binary classification) based…

↓621 year ago

ibm-research/biomed.sm.mv-te-84m-MoleculeNet-ligand_scaffold-LIPOPHILICITY-101

by ibm-research

# ibm/biomed.sm.mv-te-84m-MoleculeNet-ligand_scaffold-LIPOPHILICITY-101 biomed.sm.mv-te-84m is a multimodal biomedical foundation model for small molecules created using MMELON (Multi-view Molecular Embedding with Late Fusion), a flexible approach to aggregate multiple views (sequence, image,…

↓1.8K1 year ago

ibm-research/biomed.sm.mv-te-84m-MoleculeNet-ligand_scaffold-SIDER-101

by ibm-research

# ibm/biomed.sm.mv-te-84m-MoleculeNet-ligand_scaffold-SIDER-101 biomed.sm.mv-te-84m is a multimodal biomedical foundation model for small molecules created using MMELON (Multi-view Molecular Embedding with Late Fusion), a flexible approach to aggregate multiple views (sequence, image, graph) of…

↓101 year ago

ibm-research/biomed.sm.mv-te-84m-MoleculeNet-ligand_scaffold-BACE-101

by ibm-research

# ibm/biomed.sm.mv-te-84m-MoleculeNet-ligand_scaffold-BACE-101 biomed.sm.mv-te-84m is a multimodal biomedical foundation model for small molecules created using MMELON (Multi-view Molecular Embedding with Late Fusion), a flexible approach to aggregate multiple views (sequence, image, graph) of…

↓1.6K2 months ago

ibm-research/biomed.sm.mv-te-84m

by ibm-research

# ibm-research/biomed.sm.mv-te-84m biomed.sm.mv-te-84m is a multimodal biomedical foundation model for small molecules created using MMELON (Multi-view Molecular Embedding with Late Fusion), a flexible approach to aggregate multiple views (sequence, image, graph) of molecules in a foundation model…

↓12.1K2 months ago

gbyuvd/chemembed-chemselfies

by gbyuvd

sentence-similarity

ChemFIE-BED is a sentence-transformers based on gbyuvd/chemselfies-base-bertmlm fine-tuned on around (for now) 2 million pairs of valid molecules' SELFIES (Krenn et al. 2020) taken from COCONUTDB (Sorokina et al. 2021) and ChemBL34 (Zdrazil et al. 2023).

↓6326 months ago

gbyuvd/chemselfies-base-bertmlm

by gbyuvd

This model is a lightweight model pre-trained on SELFIES (Self-Referencing Embedded Strings) representations of molecules. It is trained on 2.7M unique and valid molecules taken from COCONUTDB and ChemBL34, with 7.3M total generated masked examples.

↓137 months ago

littleworth/protgpt2-distilled-medium

by littleworth

text-generation

A compact protein language model distilled from ProtGPT2 using complementary-regularizer distillation---a method that combines uncertainty-aware position weighting with calibration-aware label smoothing to achieve 31% better perplexity than standard knowledge distillation at 3.8x compression.

↓652 months ago

littleworth/protgpt2-distilled-tiny

by littleworth

text-generation

A compact protein language model distilled from ProtGPT2 using complementary-regularizer distillation---a method that combines uncertainty-aware position weighting with calibration-aware label smoothing to achieve 87% better perplexity than standard knowledge distillation at 20x compression.

↓202 months ago

ncfrey/ChemGPT-1.2B

by ncfrey

text-generation

# ChemGPT 1.2B ChemGPT is based on the GPT-Neo model and was introduced in the paper Neural Scaling of Deep Chemical Models.

↓2.6K3 years ago

ncfrey/ChemGPT-19M

by ncfrey

text-generation

# ChemGPT 19M ChemGPT is based on the GPT-Neo model and was introduced in the paper Neural Scaling of Deep Chemical Models.

↓6.9K3 years ago

ncfrey/ChemGPT-4.7M

by ncfrey

text-generation

# ChemGPT 4.7M ChemGPT is based on the GPT-Neo model and was introduced in the paper Neural Scaling of Deep Chemical Models.

↓4.1K3 years ago

seyonec/ChemBERTa-zinc-base-v1

by seyonec

Deep learning for chemistry and materials science remains a novel field with lots of potiential. However, the popularity of transfer learning based methods in areas such as NLP and computer vision have not yet been effectively developed in computational chemistry + machine learning.

↓254.7K5 years ago

Prior-Labs/tabpfn_2_6

by Prior-Labs

tabular-classification

### Model Overview TabPFN-2.6 is a transformer-based foundation model that uses in-context-learning to solve tabular prediction problems in a forward pass. Inference code can be found at https://github.com/PriorLabs/tabPFN.

↓11.6K1 month ago

SaeedLab/ProteoRift

by SaeedLab

feature-extraction

Github | Cite

↓52 months ago

InstaDeepAI/instanovo-phospho-v1.0.0

by InstaDeepAI

text-generation

InstaNovo-P is a specialized transformer-based model for de novo peptide sequencing from phosphoproteomics mass spectrometry data. This model is specifically trained and optimized for identifying phosphorylated peptides and their modification sites.

↓282 weeks ago

InstaDeepAI/instanovo-v1.0.0

by InstaDeepAI

text-generation

# InstaNovo: De novo Peptide Sequencing Model ## Model Description

↓352 weeks ago

InstaDeepAI/instanovo-v1.1.0

by InstaDeepAI

text-generation

# InstaNovo: De novo Peptide Sequencing Model ## Model Description

↓342 weeks ago

andrewdalpino/ESM2-150M-Protein-Cellular-Component

by andrewdalpino

text-classification

An Evolutionary-scale Model (ESM) for protein function prediction from amino acid sequences using the Gene Ontology (GO). Based on the ESM2 Transformer architecture, pre-trained on UniRef50, and fine-tuned on the AmiGO dataset, this model predicts the GO subgraph for a particular protein sequence -…

↓2611 months ago

andrewdalpino/ESM2-150M-Protein-Molecular-Function

by andrewdalpino

text-classification

An Evolutionary-scale Model (ESM) for protein function prediction from amino acid sequences using the Gene Ontology (GO). Based on the ESM2 Transformer architecture, pre-trained on UniRef50, and fine-tuned on the AmiGO dataset, this model predicts the GO subgraph for a particular protein sequence -…

↓1311 months ago

andrewdalpino/ESM2-35M-Protein-Biological-Process

by andrewdalpino

text-classification

An Evolutionary-scale Model (ESM) for protein function prediction from amino acid sequences using the Gene Ontology (GO). Based on the ESM2 Transformer architecture, pre-trained on UniRef50, and fine-tuned on the AmiGO dataset, this model predicts the GO subgraph for a particular protein sequence -…

↓2811 months ago

andrewdalpino/ESM2-35M-Protein-Molecular-Function

by andrewdalpino

text-classification

An Evolutionary-scale Model (ESM) for protein function prediction from amino acid sequences using the Gene Ontology (GO). Based on the ESM2 Transformer architecture, pre-trained on UniRef50, and fine-tuned on the AmiGO dataset, this model predicts the GO subgraph for a particular protein sequence -…

↓2011 months ago

andrewdalpino/ESM2-35M-Protein-Cellular-Component

by andrewdalpino

text-classification

An Evolutionary-scale Model (ESM) for protein function prediction from amino acid sequences using the Gene Ontology (GO). Based on the ESM2 Transformer architecture, pre-trained on UniRef50, and fine-tuned on the AmiGO dataset, this model predicts the GO subgraph for a particular protein sequence -…

↓1711 months ago

songlab/gpn-brassicales

by songlab

# GPN trained on Arabidopsis thaliana and 7 other Brassicales See https://github.com/songlab-cal/gpn for more details.

↓4841 year ago

zhihan1996/DNA_bert_6

by zhihan1996

↓27.4K10 months ago

zhihan1996/DNA_bert_5

by zhihan1996

↓48510 months ago

zhihan1996/DNA_bert_4

by zhihan1996

↓46910 months ago

zhihan1996/DNA_bert_3

by zhihan1996

↓1.9K10 months ago

BGI-HangzhouAI/Genos-m

by BGI-HangzhouAI

text-generation

Genos-m is a foundation model for human-associated microbial genomes. It is trained to model microbial DNA sequences at single-nucleotide resolution and supports ultra-long genomic contexts up to one million tokens.

↓244 days ago

AIRI-Institute/moderngena-base

by AIRI-Institute

# ModernGENA base ModernGENA is a DNA foundation model based on ModernBERT (a modernized BERT-style encoder architecture) adapted for genomic sequence modeling. ModernGENA base is the 377M-parameter version introduced in the paper Back to BERT in 2026: ModernGENA as a Strong, Efficient Baseline for…

↓4951 month ago

← Prev

1
2
3
4

Submit a resource bio.tools Awesome Bioinformatics