Cell Annotation & Type Identification

A case-file guide to identifying every cellular witness before the model starts reasoning.

Last reviewed: July 2026 · Inclusion: maintained, benchmarked, or widely cited tools · Found an issue? Suggest a fix on GitHub

Healshu as Miss Marple: examining cell markers with a magnifier, matching a cell to a village photo album, and tying a name tag onto an anonymous cell
Agatha Christie lens · identify the witness

Every cell needs an identity before it can testify.

In a detective story, a clue is only useful after you know who gave it, whether the witness is reliable, and which detail is a red herring. Cell annotation works the same way: marker genes, reference atlases, probability scores, and expert review turn anonymous expression profiles into interpretable biological identities.

Healshu as Miss Marple knitting in a library armchair, a strand of yarn leading to a small covered figure with a question-mark tag on the rug

The Body in the Library

Agatha Christie's The Body in the Library (1942), re-read as a cell-annotation parable: a nameless young woman is found in the Bantrys' respectable library, and Miss Marple can only give her back a name — by reading small telltale details and village parallels. (Classic-novel spoilers inside.)

Scene 1 · The body in the library

"Who is she?"

The Bantrys find a young woman's body on their library hearthrug. She belongs to no one in the village, and nothing about her announces a name.

The annotation problem. An anonymous expression profile lands in your matrix. Before it can testify — before any downstream analysis — it needs an identity.

Scene 2 · Bitten nails and cheap dresses

Small details betray a name

Miss Marple doubts the official identification because of the girl's fingernails, her teeth, her clothes. Identity hides in details others dismiss as trivial.

Marker genes. A handful of telltale signals gives a cell away. Markers are the bitten nails of a dataset — small, specific, and easy to overlook.

Scene 3 · St. Mary Mead parallels

"She reminds me of old Mr. Harbottle"

Miss Marple solves cases by analogy: every stranger resembles someone she has known for decades in her village. Her village is a living reference atlas.

Reference atlases. Azimuth, CellTypist, SingleR and Pan-human Azimuth do the same: match an unknown cell against curated references — "this one looks like a T cell I know".

Scene 4 · The wrong name on the tag

First label: Ruby Keene. Truth: Pamela Reeves

The body is confidently identified as a missing dancer — and the identification is wrong. The wrong name nearly sends the whole investigation astray.

Misannotation. A confident first label is not the truth. Keep confidence scores visible, and let low-evidence cells stay unknown rather than forcing a name.

Scene 5 · Gossip is evidence too

Every rumor gets cross-examined

Miss Marple listens to all the village gossip — then checks each piece against independent facts before believing any of it.

Validation. Corroborate labels with independent evidence: confidence distributions, marker re-expression, and reference databases (CellMarker 2.0, PanglaoDB, HuBMAP).

Epilogue · Every body has a name

The name, with its evidence attached

By the end, the nameless girl has her real name back — and every step of the identification can be retraced, clue by clue.

The annotated matrix. The goal is not just labels, but labels with provenance: type, confidence, and the evidence chain that earned them.

Four-panel comic retelling The Body in the Library as a cell-annotation parable: an unidentified cell is found, its marker genes are read as telltale details, it is matched against a reference atlas of known cells, and it is finally given a name with its confidence recorded
Four scenes, one identification: the anonymous cell, the telltale markers, the village parallel, and the name that finally sticks — with its evidence attached.

In St. Mary Mead, nobody is anonymous. That is the whole point of annotation: there are no anonymous cells — only cells whose evidence you have not read yet.

Contents

Overview

Cell annotation assigns biological identities (cell types, states) to individual cells based on their transcriptomic profiles. This page provides a comprehensive guide to available tools, helping you choose the right method for your analysis.

UMAP comparison: anonymous grey question-mark cells on the left become labeled, color-coded cell types on the right, with one low-evidence cluster honestly left unassigned
From anonymous profiles to identified witnesses: annotation assigns interpretable identities — and leaves low-evidence cells unassigned rather than forcing a label.
case file witness lineup

St. Mary Mead logic

When many witnesses are plausible, compare multiple evidence sources: reference labels, marker genes, cluster context, and confidence scores.

pattern or red herring?

Bitten nails & red herrings

A strong marker pattern can be a clue, but the method still has to test whether it is biology, batch structure, or an annotation artifact.

eliminate impossible labels

The wrong name on the tag

Good annotation is often elimination: reject impossible labels, keep uncertainty visible, then validate the final identity with independent evidence.

General Annotation Workflow

Preprocess

QC & normalize

Select Tool

Based on data & needs

Annotate

Run classification

Validate

Check markers

Unknown cell Evidence Suspects Label + confidence
Annotation is a chain of evidence: expression profile -> marker/reference evidence -> candidate labels -> final identity with confidence and uncertainty.

Validate the verdict

Confidence spectrum from 0 to 1: low-confidence question-mark cells are set aside into an unknown basket, while high-confidence cells receive labels such as T cell, B cell and Mono
  • Score the confidence, not just the label. Inspect per-cell and per-cluster confidence distributions; keep low-confidence cells as unknown instead of forcing a verdict.
  • Re-check canonical markers. Confirm that assigned types re-express their expected marker genes in your own data.
  • Cross-examine with reference databases: CellMarker 2.0 PanglaoDB Azimuth · HuBMAP atlases Human Cell Atlas

Quick Reference

Subway map of all 21 cell annotation tools: five colored lines for reference-based, marker-based, machine learning, foundation model, and specialized methods, with each tool as a station
All 21 tools on one map, grouped by methodological concept — curated family assignments and full details in the table below.
Tool Family Year Approach Best For Links
scmap Reference 2018 Projection to reference Fast large-scale annotation Paper | GitHub
SingleR Reference 2019 Correlation-based Bulk reference compatibility Paper | GitHub
CellAssign Marker genes 2019 Probabilistic marker-based Multi-batch studies with markers Paper | GitHub
SCINA Marker genes 2019 Semi-supervised mixture Simple marker-based Paper | GitHub
SingleCellNet ML classifier 2019 Random forest Cross-platform robustness Paper | GitHub
CHETAH Marker genes 2019 Hierarchical tree Uncertain/intermediate cells Paper | GitHub
CellTypist ML classifier 2022 Logistic regression Immune cell typing Paper | GitHub
Azimuth Reference 2023 Reference mapping (Seurat) Seurat ecosystem users Paper | GitHub
Pan-human Azimuth Reference 2026 Supervised NN, organism-scale typology Organism-wide reference mapping Paper
scGPT Foundation 2024 Foundation model (Transformer) Multi-task learning Paper | GitHub
scFoundation Foundation 2024 Foundation model (100M params) Low-depth data enhancement Paper | GitHub
CytoTRACE 2 Specialized 2025 Interpretable deep learning Developmental potential Paper | GitHub
scCATCH Marker genes 2020 Cluster-based + database Automatic cluster annotation Paper | GitHub
scDeepSort ML classifier 2021 Pre-trained GNN Reference-free annotation Paper | GitHub
scBERT ML classifier 2022 BERT-based transformer Deep learning annotation Paper | GitHub
Concerto ML classifier 2022 Contrastive learning Million-scale mapping Paper | GitHub
CellHint ML classifier 2023 Harmonization (PCT) Cross-dataset standardization Paper | GitHub
UCE Foundation 2023 Universal embeddings Zero-shot cross-species Paper | GitHub
scInterpreter Foundation 2024 LLM-based (Llama) Gene knowledge integration Paper (arXiv)
TCellSI Specialized 2024 T cell state scoring 8 T cell functional states Paper | GitHub
AIDO.Cell Foundation 2024 Full-transcriptome transformer 19K gene context Paper | GitHub
Nicheformer Foundation 2025 Spatial + dissociated FM Spatial context prediction Paper | GitHub

Reference-Based Methods

Map query cells to pre-annotated reference datasets

Four-panel comic of reference-based annotation with Azimuth: an unknown cell arrives without a name, it is held up against a curated reference atlas of known cell types, the closest match transfers its label along with a mapping score, and cells that match nothing well are left flagged rather than forced into a label
Reference mapping is Miss Marple’s village method in software: the atlas supplies the people you already know, and a query cell earns a name only when the resemblance — and its score — holds up.
scmap Best for: Fast large-scale annotation 2018
Nature Methods | Wellcome Sanger Institute
Fast Projection Bioconductor
First scalable projection method for single-cell reference mapping. Offers two modes: scmap-cluster (fast, maps to cluster centroids) and scmap-cell (precise, maps to individual cells).
  • Dropout-based feature selection
  • Handles "unassigned" cells gracefully
  • Low memory footprint
  • Requires pre-computed indices for references
SingleR Best for: Bulk reference compatibility 2019
Nature Immunology | Weizmann Institute
Bulk Compatible Correlation Pre-built Refs
Correlates single-cell profiles with reference datasets (bulk or single-cell). Iteratively refines annotations using variable genes between top candidates.
  • Works with bulk RNA-seq references
  • Pre-built references (HPCA, Blueprint, Monaco)
  • Iterative fine-tuning refinement
  • Can be slow on very large datasets
Azimuth Best for: Seurat ecosystem users 2023
Cell (Seurat v4) | Satija Lab
Web App Seurat Multi-Modal ATAC-seq
Automated reference-based annotation using HuBMAP atlases. Single-command RunAzimuth() provides multi-resolution labels (l1, l2) with confidence scores. Supports ATAC-seq via bridge integration.
  • No-code web interface available
  • HuBMAP reference atlases (PBMC, lung, kidney...)
  • Bridge integration for ATAC-seq
  • Restricted to provided HuBMAP references
Pan-human Azimuth Best for: Organism-wide reference mapping 2026
bioRxiv | NYGC · Satija Lab
Organism-Scale Hierarchical Typology HuBMAP Spatial-Ready
Supervised neural network that maps human cells from diverse tissues and datasets onto a single hierarchical organism-scale typology — the whole-organism successor of Azimuth, developed through NIH HuBMAP with a uniformly curated, QC-stringent training corpus.
  • One unified reference across tissues and technologies
  • Mapped tens of millions of cells (Tabula Sapiens, scBaseCamp)
  • Extends to spatial transcriptomics: recovers kidney cortical structures, glomerular states matching expert pathology
  • Cloud, R, and Python interfaces
  • bioRxiv preprint (July 2026); peer review pending

Marker Gene-Based Methods

Leverage prior knowledge of cell type marker genes

CellAssign Best for: Multi-batch studies with markers 2019
Nature Methods | BC Cancer
Batch Correction Probabilistic TensorFlow
Probabilistic framework using marker genes with explicit batch effect modeling. Provides uncertainty quantification and robust to ~30% marker misspecification.
  • Hierarchical Bayesian model (negative binomial)
  • Explicit batch/patient/sample modeling
  • GPU acceleration via TensorFlow
  • Performance sensitive to marker list quality/specificity
SCINA Best for: Simple marker-based annotation 2019
Genes | UT Southwestern
Simple Semi-Supervised EM Algorithm
Semi-supervised annotation using bimodal Gaussian mixture models for marker genes. Simple input: just provide marker gene lists per cell type.
  • Bimodal on/off expression model
  • No reference dataset required
  • Easy to use - just marker lists
  • Relies heavily on assumption of bimodal gene expression
CHETAH Best for: Uncertain/intermediate cells 2019
Nucleic Acids Research
Hierarchical Uncertainty Tree-Based
Hierarchical classification using reference tree structure. Stops at appropriate resolution when confidence drops, flagging uncertain cells.
  • Auto-determines annotation granularity
  • Identifies intermediate/ambiguous states
  • Confidence thresholds at each tree level
  • High rate of "unassigned" cells if reference is biologically distinct
scCATCH Best for: Automatic cluster annotation 2020
iScience | Zhejiang University
Cluster-Based CellMatch DB Evidence-Based
Automatic cluster annotation using CellMatch database (353 cell types, 20,792 markers across 184 tissues). Evidence-based scoring ranks candidates by marker matches and literature support.
  • Paired cluster comparison reduces false positives
  • CellMatch integrates CellMarker, MCA, CancerSEA
  • 83% average accuracy across tissues
  • Database-dependent; novel cell types may not be recognized

Machine Learning Classifiers

Traditional ML approaches for supervised classification

SingleCellNet Best for: Cross-platform robustness 2019
Cell Systems | Johns Hopkins
Cross-Platform Random Forest Calibrated
Random forest classifier using "top-pair" gene features that are robust to technical variation. Provides calibrated scores for quality assessment.
  • Rank-based gene pairs for cross-platform robustness
  • Train custom classifiers from your reference
  • Calibrated probability scores
  • Requires training data that closely matches target data type
CellTypist Best for: Immune cell typing 2022
Science | Wellcome Sanger Institute
Immune Cells Cross-Tissue 360K Cells 101 Types
Fast logistic regression trained on cross-tissue immune atlas (360K cells, 16 tissues, 101 cell types). Continuously updatable as new data becomes available.
  • F1: 0.95 (high) / 0.89 (low hierarchy)
  • Hierarchical: 32 broad → 91 fine types
  • SGD training - fast and scalable
  • Currently optimized primarily for immune cells
scDeepSort Best for: Reference-free annotation 2021
Nucleic Acids Research | Zhejiang University
Pre-trained GNN 265K Cells
First pre-trained graph neural network for cell type annotation. Treats cells and genes as graph nodes with expression as weighted edges. No additional reference required at prediction time.
  • 83.79% accuracy across 265,489 cells
  • Weighted graph aggregator handles batch effects
  • Outperforms 12 existing methods
  • Pre-trained on specific cell types; may struggle with novel types
scBERT Best for: Deep learning annotation 2022
Nature Machine Intelligence | Tencent
BERT Transformer Performer
BERT-based model adapted for scRNA-seq with Performer encoder. Two-stage pre-training + fine-tuning paradigm. Attention weights enable discovery of cell-type-specific genes.
  • Strong batch effect resistance
  • Interpretable attention for gene discovery
  • Novel cell type detection capability
  • Requires GPU for efficient training/inference
Concerto Best for: Million-scale mapping 2022
Nature Machine Intelligence | BGI
Contrastive Learning Self-Distillation Multi-Modal
Contrastive learning framework using asymmetric teacher-student self-distillation. First application of contrastive learning to single-cell. Supports multimodal RNA + protein integration.
  • Rapid mapping to million-scale atlases
  • NOTA (None-of-the-above) rejection for novel types
  • Hierarchical fine-grained annotation
  • Requires large training datasets for optimal performance
CellHint Best for: Cross-dataset standardization 2023
Cell | Wellcome Sanger Institute
Harmonization HCA PCT
Predictive clustering tree (PCT) tool for harmonizing cell type annotations across datasets. Creates hierarchical cell type relationships for standardized Human Cell Atlas integration.
  • Overcomes batch effects in cross-dataset comparisons
  • Annotation-aware data integration
  • From CellTypist team - compatible ecosystem
  • Requires multiple datasets with existing annotations

Foundation Models

Large-scale pre-trained models for general single-cell analysis

scGPT Best for: Multi-task learning 2024
Nature Methods | University of Toronto
Multi-Task Transformer 33M Cells Fine-Tunable
Foundation model trained on 33M human cells. Simultaneously learns cell and gene representations. Fine-tune for annotation, perturbation prediction, batch integration, and more.
  • Transformer with specialized attention masks
  • <cls> token for cell-level representation
  • Gene network inference from attention
  • Requires significant GPU resources for fine-tuning
scFoundation Best for: Low-depth data enhancement 2024
Nature Methods | BioMap
100M Params 50M Cells Low-Depth xTrimoGene
100M parameter model with xTrimoGene architecture and Read-Depth Aware (RDA) pre-training. Excels at gene expression enhancement and handles variable sequencing depths.
  • Asymmetric encoder-decoder (non-zero only)
  • RDA task handles low-depth data
  • Zero-shot expression enhancement
  • Very large model size makes local deployment challenging
UCE Best for: Zero-shot cross-species 2023
bioRxiv | Stanford & CZI
Cross-Species 650M Params 36M Cells ESM2
Universal Cell Embeddings using "Bags of RNA" approach with ESM2 protein embeddings. Zero-shot cell type classification across species without homolog mapping. 33-layer transformer on Integrated Mega-scale Atlas.
  • No cell type annotations needed for training
  • Cross-species annotation without fine-tuning
  • 1280-dimensional universal embeddings
  • Currently transcriptomics only; preprint status
scInterpreter Best for: Gene knowledge integration 2024
arXiv Preprint | Chinese Academy of Sciences
LLM-Based Llama-13b GPT Embeddings Preprint
Adapts LLMs (Llama-13b) to interpret scRNA-seq using GPT-3.5 gene description embeddings from NCBI. Bridges biological knowledge from language models with expression profiles.
  • LLM-encoded biological knowledge integration
  • Frozen LLM + lightweight projection layers
  • Gene-level semantic grounding via NCBI
  • Preprint; code not publicly available yet
AIDO.Cell Best for: Full 19K gene context 2024
NeurIPS | GenBio AI
Full Transcriptome 650M Params FlashAttention 19K Context
First single-cell FM to process entire 20K-gene transcriptome without truncation using dense Transformer + FlashAttention-2. Scaling study from 3M to 650M params on 50M cells.
  • No gene truncation/sampling - full context
  • Auto-discretization learns flexible expression embeddings
  • Read-depth aware pre-training
  • Requires 256 H100 GPUs for training; inference still heavy
Nicheformer Best for: Spatial context prediction 2025
Nature Methods | Helmholtz Munich
Spatial + Dissociated 110M Cells Multi-Tech
First foundation model trained on both dissociated (57M) and spatial (53M) transcriptomics. Learns spatially-aware representations enabling transfer of spatial context to scRNA-seq.
  • Predicts cellular microenvironments from expression
  • SpatialCorpus-110M across multiple technologies
  • Transfers spatial annotations to dissociated data
  • Spatial predictions may not generalize to all tissue contexts

Developmental Potential & Specialized

Methods for predicting cell potency and developmental states

CytoTRACE 2 Best for: Developmental potential 2025
Nature Methods | Stanford
Interpretable Potency Cross-Dataset GSBN
Predicts absolute developmental potential (0-1 scale) and potency categories (totipotent → differentiated) using interpretable Gene Set Binary Networks. Works across datasets without batch correction.
  • 6 potency categories (totipotent to differentiated)
  • Binary weights enable gene set extraction
  • τ = 0.82 correlation with ground truth
  • Focuses on potency state, not specific cell type identity
TCellSI Best for: T cell functional states 2024
iMeta | Huazhong University of Science and Technology
T Cells State Inference 8 States R Package
Specialized tool for inferring 8 distinct T cell functional states from scRNA-seq: Quiescence, Regulating, Proliferation, Helper, Cytotoxicity, Progenitor exhaustion, Terminal exhaustion, and Senescence.
  • 8 curated T cell functional state signatures
  • Mann-Whitney U statistics for robust scoring
  • Works with bulk RNA-seq and scRNA-seq
  • Specific to T cells; not generalizable to other cell types

Which Tool Should I Use?

Start here Trusted reference atlas? Azimuth · CellTypist · SingleR · scmap Only marker genes? CellAssign · SCINA · CHETAH Nothing curated yet? scGPT · UCE · scFoundation Special states? CytoTRACE 2 · TCellSI · Concerto

Have a Reference Atlas?

  • Seurat user: Azimuth
  • Organism-wide reference: Pan-human Azimuth
  • Immune cells: CellTypist
  • Bulk reference: SingleR
  • Fast/simple: scmap

Have Marker Genes?

  • Multi-batch data: CellAssign
  • Simple/quick: SCINA
  • Uncertain cells: CHETAH

Multiple Tasks / Foundation Models?

  • Annotation + Integration + Perturbation: scGPT
  • Low-depth data: scFoundation
  • Cross-species zero-shot: UCE
  • Full 19K gene context: AIDO.Cell
  • Spatial + scRNA-seq: Nicheformer

Specialized Cell Types/States?

  • Developmental potency: CytoTRACE 2
  • T cell functional states: TCellSI
  • Cross-dataset harmonization: CellHint
  • Million-scale mapping: Concerto

🛠️ Hands-On Practice

The walkthrough below takes a QC'd AnnData object and gives every cell a name, end to end: normalize to the exact scale CellTypist expects, run a pretrained immune model with majority voting over Leiden clusters, keep the per-cell probability as a confidence score, cross-examine the labels against a hand-picked marker panel, and leave low-evidence cells honestly marked Unknown. Everything is computational and retrospective — the only evidence used is the count matrix you already have.

Environment & packages

CellTypist is a light logistic-regression classifier and installs cleanly alongside Scanpy. leidenalg supplies the over-clustering used for majority voting; decoupler is optional but convenient for signature-based marker scoring. Keep an R environment on hand only if you want the SingleR cross-check.

# conda / mamba recommended
conda create -n scannot python=3.10 -y
conda activate scannot

pip install scanpy celltypist leidenalg
pip install decoupler            # optional: signature / marker scoring

# optional R cross-check with SingleR
# R: BiocManager::install(c("SingleR", "celldex", "SingleCellExperiment"))

Hardware. CellTypist inference is CPU-only and fast: a 50k-cell dataset annotates in under a minute on a laptop, and 500k cells fit comfortably on a 32–64 GB compute node. Only the foundation-model route (scGPT, UCE, scFoundation fine-tuning) needs a GPU; reference mapping with Azimuth or scArches sits in between.

Data structures & formats

  • AnnData (input) — must carry log1p-normalized expression in adata.X; stash the raw integers in adata.layers["counts"] before normalizing so differential expression and doublet tools can still reach them
  • CellTypist model (.pkl) — a pickled logistic-regression classifier plus its label set; downloaded once into ~/.celltypist/data/models/. Immune_All_Low.pkl covers fine-grained immune subtypes, Immune_All_High.pkl the coarse ones
  • AnnotationResult — returned by celltypist.annotate(), with .predicted_labels (a DataFrame holding predicted_labels, over_clustering, majority_voting), .probability_matrix (cells × labels), and .decision_matrix (raw decision scores)
  • Per-cell obs columns — after result.to_adata(): adata.obs["predicted_labels"], adata.obs["majority_voting"], adata.obs["conf_score"]
  • Marker dictionary — a plain dict of {"cell type": [genes]} fed to sc.pl.dotplot(); the orthogonal evidence you judge the classifier against
  • h5ad — write the annotated object with adata.write_h5ad(); for the R/SingleR route export counts + obs to a SingleCellExperiment or via zellkonverter::readH5AD()

Minimal code walkthrough

Load a QC-filtered object, normalize to exactly 10,000 counts per cell and log1p-transform, build a Leiden over-clustering, run CellTypist with majority_voting=True, record per-cell confidence, demote low-confidence cells to Unknown, and cross-check the result with a marker dot plot.

import scanpy as sc
import numpy as np
import celltypist
from celltypist import models

# 1. Load the QC'd object (post filtering / doublet removal, raw counts in X)
adata = sc.read_h5ad("adata_qc.h5ad")
adata.layers["counts"] = adata.X.copy()      # keep integers for later DE

# 2. NORMALIZATION CONTRACT — the classic CellTypist mistake.
#    CellTypist models were trained on expression normalized to exactly
#    10,000 counts per cell and then log1p-transformed. target_sum=1e4 is
#    NOT optional: use 1e6 (CPM), SCTransform residuals, or scaled/z-scored
#    values and the probabilities are silently wrong, not an error.
sc.pp.normalize_total(adata, target_sum=1e4)
sc.pp.log1p(adata)
adata.raw = adata                            # log-normalized snapshot

# 3. Over-clustering for majority voting. CellTypist can compute its own,
#    but reusing your analysis Leiden keeps labels consistent with the UMAP
#    you will actually show. Slightly over-cluster (higher resolution).
sc.pp.highly_variable_genes(adata, n_top_genes=2000, batch_key=None)
sc.pp.pca(adata, n_comps=50, mask_var="highly_variable")
sc.pp.neighbors(adata, n_neighbors=15)
sc.tl.leiden(adata, resolution=1.5, key_added="leiden", flavor="igraph",
             n_iterations=2, directed=False)
sc.tl.umap(adata)

# 4. Download and load a pretrained model. Immune_All_Low = fine-grained
#    immune subtypes; Immune_All_High = coarse lineages. Pick the resolution
#    your tissue actually supports.
models.download_models(model=["Immune_All_Low.pkl"])
model = models.Model.load(model="Immune_All_Low.pkl")
print(model.cell_types[:10], "...", len(model.cell_types), "labels")

# 5. Annotate. majority_voting=True refines per-cell calls by taking the
#    dominant label within each over-clustering group -- it removes salt-and-
#    pepper noise but can erase genuinely rare populations (see pitfalls).
predictions = celltypist.annotate(
    adata,
    model=model,
    majority_voting=True,
    over_clustering="leiden",
)

# 6. Per-cell confidence = max class probability. Keep it visible; it is the
#    only honest signal that a cell was a coin flip between two labels.
prob = predictions.probability_matrix
adata.obs["predicted_labels"]  = predictions.predicted_labels["predicted_labels"].values
adata.obs["majority_voting"]   = predictions.predicted_labels["majority_voting"].values
adata.obs["conf_score"]        = prob.max(axis=1).values
adata.obs["runner_up"]         = prob.apply(
    lambda r: r.sort_values(ascending=False).index[1], axis=1
).values

# 7. Do NOT force a name on weak evidence. Demote low-probability cells to
#    "Unknown" rather than letting the argmax invent an identity.
LOW_CONF = 0.5
adata.obs["cell_type"] = adata.obs["majority_voting"].astype(str)
adata.obs.loc[adata.obs["conf_score"] < LOW_CONF, "cell_type"] = "Unknown"
n_unknown = (adata.obs["cell_type"] == "Unknown").sum()
print(f"Unknown (conf < {LOW_CONF}): {n_unknown} / {adata.n_obs} cells")
print(adata.obs["cell_type"].value_counts().head(20))

# 8. ORTHOGONAL CROSS-EXAMINATION. A classifier label is a hypothesis; a
#    canonical marker panel is the independent witness. Every called type
#    should light up its own markers and stay dark elsewhere.
markers = {
    "T cell":        ["CD3D", "CD3E", "TRAC"],
    "CD8 T":         ["CD8A", "CD8B", "GZMK"],
    "CD4 T":         ["IL7R", "CCR7", "CD40LG"],
    "NK":            ["NKG7", "GNLY", "KLRD1"],
    "B cell":        ["MS4A1", "CD79A", "CD79B"],
    "Plasma":        ["JCHAIN", "MZB1", "XBP1"],
    "Monocyte":      ["LYZ", "CD14", "FCN1"],
    "cDC":           ["FCER1A", "CD1C", "CLEC9A"],
    "Platelet":      ["PPBP", "PF4", "ITGA2B"],
}
markers = {k: [g for g in v if g in adata.var_names] for k, v in markers.items()}
sc.pl.dotplot(adata, markers, groupby="cell_type",
              standard_scale="var", dendrogram=True)
sc.pl.umap(adata, color=["cell_type", "conf_score"], legend_loc="on data")

# 9. Quantify agreement between the two lines of evidence: score each cell
#    for each marker panel and check the argmax against the classifier.
for ct, genes in markers.items():
    sc.tl.score_genes(adata, genes, score_name=f"score_{ct}")
scores = adata.obs[[f"score_{ct}" for ct in markers]]
adata.obs["marker_call"] = (
    scores.idxmax(axis=1).str.replace("score_", "", regex=False)
)
print(sc.metrics.confusion_matrix("cell_type", "marker_call", adata.obs))

# 10. Save. Store the full probability matrix so downstream reviewers can
#     re-threshold without re-running the classifier.
adata.obsm["celltypist_prob"] = prob.reindex(adata.obs_names).values
adata.uns["celltypist_labels"] = list(prob.columns)
adata.write_h5ad("adata_annotated.h5ad")
print("Saved adata_annotated.h5ad")

Fallback routes. If no CellTypist model fits your tissue, two alternatives use the same annotated object. SingleR (R) correlates each cell against a bulk or single-cell reference and reports a per-cell delta score, so it flags weak assignments explicitly:

# R -- SingleR cross-check on the same log-normalized matrix
library(SingleR); library(celldex); library(zellkonverter)

sce <- readH5AD("adata_annotated.h5ad")            # X = log-normalized
assayNames(sce) <- "logcounts"
ref <- celldex::HumanPrimaryCellAtlasData()        # or MonacoImmuneData()

pred <- SingleR(test = sce, ref = ref, labels = ref$label.fine,
                de.method = "wilcox")
table(pred$pruned.labels, useNA = "ifany")         # NA = pruned, low confidence
plotDeltaDistribution(pred)                        # per-cell assignment margin

Or take the reference-mapping route — Azimuth (Seurat) or scArches/scANVI — which projects your query into a curated atlas embedding and transfers labels with a mapping score. That route handles batch effects between query and reference far better than a plain classifier, at the cost of needing a well-matched atlas.

Common pitfalls & tips

  • The 10,000-count normalization is a hard requirement. CellTypist expects sc.pp.normalize_total(adata, target_sum=1e4) followed by sc.pp.log1p(). Feeding CPM (1e6), raw counts, scaled/z-scored values, or SCTransform residuals produces confidently wrong labels with no warning — the model still returns probabilities, they are just meaningless. Check np.expm1(adata.X[:5]).sum(1) ≈ 10,000 before you trust anything.
  • Majority voting over-merges rare types. Smoothing to the dominant label inside each over-cluster cleans up noise, but a 30-cell pDC or MAIT population sitting inside a larger cluster is erased outright. Always keep the raw predicted_labels column beside majority_voting, and over-cluster (Leiden resolution 1.5–2.0) so rare types get their own group before voting.
  • A missing cell type gets mislabeled, not flagged. Classifiers are closed-world: if your tumor sample contains malignant epithelium and the reference only knows immune cells, those cells are assigned to the nearest immune label with a plausible-looking probability. Absence of a class is never reported as absence — inspect the confidence distribution and the markers of every cluster, especially large ones with a suspiciously uniform label.
  • Tissue-mismatched references are a silent failure. Immune_All_Low on brain tissue will call microglia something in its immune vocabulary; a gut atlas will not describe kidney. Match the reference to the tissue and the assay (10x vs Smart-seq2 differ in gene detection), and prefer the coarse model when the fine one has no vocabulary for your sample.
  • Batch effects break reference mapping, not just integration. Query-vs-reference differences in chemistry, donor, or dissociation protocol shift the embedding and drag labels with them. Use a method that models the batch (scArches/scANVI, Azimuth's anchor transfer, Symphony) rather than nearest-neighbor transfer on unintegrated PCA, and never integrate query and reference so aggressively that the biological signal you are trying to read is also removed.
  • Always validate with orthogonal markers. Treat the classifier output as a hypothesis and the canonical marker dot plot as the independent witness. A "CD8 T cell" cluster with no CD3D and abundant LYZ is a misannotation regardless of how high the probability is. Disagreement between the two is the signal worth investigating, not a nuisance.
  • Mark low-confidence cells "Unknown" rather than force-assigning. The argmax always returns something. Thresholding on max probability (≈0.5, calibrated by looking at the score histogram) and labeling the rest Unknown is honest and downstream-safe; a forced label propagates into every DE test, composition comparison, and figure that follows.
  • Annotate on log-normalized data, test on counts. Keep the raw integers in adata.layers["counts"]. Annotation needs the normalized matrix, but differential expression, pseudobulk, and doublet re-checks need counts — losing them means re-running the pipeline from the start.