Skip to content

Getting started

Install

git clone https://github.com/JaneliaSciComp/chunkmirage
cd chunkmirage
uv sync --extra all --group dev        # or: pip install -e ".[all]"; add --extra gpu (or cpu) for register://

Requires Python 3.11+. Core dependencies are tensorstore, numpy, numcodecs, starlette, uvicorn, pydantic and typer. Optional extras: hdf5 (h5py), tiff (tifffile and imagecodecs, for GeoTIFFs), ops (scipy filters), mcp (planned MCP server), and for register:// either gpu (PyTorch with CUDA, about 3 GB) or cpu (its CPU-only build, for machines without an NVIDIA GPU). Neither is in all, and they cannot be installed together.

Serve something

chunkmirage serve /path/to/data.zarr/em/fibsem-uint8 --op threshold:low=120 --port 8000

The unprocessed source is served alongside as raw and shows as a second layer, at no extra cost (--no-raw to skip it).

SOURCE can be a zarr v2/v3 array or multiscale group (s0, s1, ...), an N5 dataset or group, a Neuroglancer precomputed volume, or file.h5::/dataset. Local paths and s3://, gs://, http(s):// URLs all work. The command prints:

source:       zarr3://http://localhost:8000/data/@<digest>/zarr3
neuroglancer: https://neuroglancer-demo.appspot.com/#!...
control API:  http://localhost:8000/api/datasets/data

URLs use the machine's network address so they work from other machines too. Open the Neuroglancer link in Chrome or Firefox. See FAQ: does this work with the hosted Neuroglancer?

The demo

No data needed: generate a virtual 4096³ volume (69 gigavoxels, nothing on disk) and run a small segmentation pipeline on it:

chunkmirage serve "synthetic://blobs+noise?shape=4096,4096,4096" \
    --op gaussian:sigma=1.5 --op threshold:low=110 \
    --op morphology:operation=open,radius=2 --op label:min_size=200 \
    --python-viewer

Open the printed control page: raw shows the generated volume, processed the labelled objects. Drag low and watch objects appear and merge; change radius to remove specks. The smaller disk-based demo is:

uv run python examples/demo.py --port 8000

Generates a synthetic multiscale volume of blobs, then serves it twice: raw as zarr v3 and thresh (gaussian then threshold) as a precomputed segmentation overlay. Both share one cache, so the raw chunks are read once.

Change the pipeline interactively

Open http://localhost:8000/ui for sliders generated from each op's parameters. For live updates that keep your camera position, start the server with --python-viewer and open the printed viewer URL. Details and trade-offs: Interactivity.

Change the pipeline live from the shell

curl -X PUT localhost:8000/api/datasets/thresh -H 'content-type: application/json' \
  -d '{"source": "examples/demo-data.zarr/blobs",
       "ops": [{"op": "gaussian", "sigma": 1}, {"op": "threshold", "low": 160}]}'

The response contains new source URLs with an updated digest. Viewers cache chunks by URL, so paste the new URL into the layer's source field and it refetches. Only the changed stage recomputes; see Caching.

Use it as a library

from chunkmirage import Pipeline, open_source, create_app
from chunkmirage.ops import Threshold
import uvicorn

src = open_source("s3://bucket/data.zarr/em")           # multiscale group or single array
pipe = Pipeline(src, [Threshold(low=120)])
app = create_app({"em-thresh": pipe})                   # Starlette ASGI app
uvicorn.run(app, port=8000)

Read the result from Python instead of a viewer

Anything that reads zarr over HTTP can consume a served pipeline as a lazy array:

import tensorstore as ts
arr = ts.open({"driver": "zarr3",
               "kvstore": "http://localhost:8000/em-thresh/zarr3/s0/"}).result()
arr[0:64, 0:64, 0:64].read().result()