Python API¶
Top level¶
A source's levels through ops. Ops run on every level, except from the first op
with an input voxel size (Op.input_voxel_size, a model trained at one resolution):
that one runs on one level (the source's at that size, else one resampled to it), its
output is cached, and the coarser levels are made from it by downsample, by the
source pyramid's own factors; the ops after it run on every level.
Bases: BaseModel
The query of a register:// URL. The one definition of it: the browser engine
(web/) generates its TypeScript types and form defaults from this model's JSON
Schema (chunkmirage schema).
One schema whose $defs hold every model: PipelineSpec, RegisterParams,
StitchParams, one per registered op (tagged with its op name), and OpSpec, the union of them.
Open any supported dataset as a MultiscaleSource.
Dispatch is by content, not extension: zarr v2/v3, n5 and neuroglancer precomputed go
through tensorstore; .h5/.hdf5 paths (file.h5::/dataset) go through h5py;
synthetic://kind?shape=... generates data procedurally (see sources.synthetic);
scene://group?image=...&target=... resamples an image through OME-Zarr 0.6
coordinate transformations (see sources.scene); warp://image?field=swirl&...
twists an image through a procedural deformation (see sources.warp);
register://moving?fixed=... registers one image onto another, solving the
deformation on a GPU when opened (see sources.register); stack://a|b serves
images on one grid as the channels of one array, for ops over several images (see
sources.stack); flip://image?axes=y mirrors an image stored the other way round
(see sources.flip); stitch://project.xml stitches a BigStitcher project's tiles by
interest points and RANSAC when opened, and fuses them as they are read (see
sources.stitch); .tif/.tiff paths are (cloud-optimized) GeoTIFFs, read
tile by tile with their overviews as levels (see sources.geotiff). Other packages
add schemes of their own (register_source, or the chunkmirage.sources entry
point).
cache_bytes is tensorstore's pool of decoded source chunks; cache is the chunk
cache that computed sources keep their expensive intermediates in (register://'s
refined blocks), normally the pipeline's. kind says what the values are (image,
label, mask) over what the source guessed: stored arrays guess from their dtype
(core.kind_for_dtype), so a uint32 image needs kind="image".
The ASGI app serving datasets (a mapping of names to pipelines or specs, or a
:class:DatasetRegistry) through every frontend, with the control API.
public_url: base of the links the app hands out (default: the request's own, with the prefix the app is mounted under).allow_edit: whether/api/datasetsmay create, replace and remove datasets.threads: chunk requests computing at once (default 40).token: required on/api/*(header or?token=); the datasets stay open.extra_routes: Starlette routes of your own, matched after the built-in ones and before the datasets (/{name}/...); under/api/they share the token.route_plugins: also add the routes of installedchunkmirage.routesplugins.resolver: builds datasets requested by a name nobody registered (DatasetRegistry.resolve); the registry's own if not given.
It can be mounted in another Starlette app (Mount("/prefix", app)): paths, the token
and links are then relative to the prefix.
Serve app (create_app's, or any ASGI app) on host: on port, 0 for any
free one, None for the first free one from 8000. The port is bound before this returns
or calls anything, so no other process can take it. on_ready(server) is called once
connections are accepted. By default the server runs in a background thread and is
returned running (server.port, server.url, server.stop()); with block
it runs in this thread until stopped, as chunkmirage serve does.
app served by uvicorn on sock, a socket already bound and listening
(netutil.bind_socket), so port and url are known before it runs.
run serves in this thread until stopped (in the main thread, Ctrl-C stops it);
start serves in a background thread and returns once connections are accepted.
on_ready(server) is called then, from the server's event loop, so it should be
quick. ssl is (certfile, keyfile).
url
property
¶
Where other machines reach it: this machine's network address when bound to
every interface, localhost for a loopback address.
run()
¶
Serve in this thread until stop (or Ctrl-C, in the main thread).
start(timeout=60)
¶
Serve in a background thread; returns once it accepts connections.
stop(timeout=10)
¶
Stop serving: requests in progress finish, then the port is released.
wait()
¶
Block until a server started in the background stops.
Named pipelines sharing one cache. Thread-safe for live edits.
version increments on every add/remove; subscribe registers a callback invoked
(outside the lock, in the editing thread) with (name, pipeline_or_None).
resolver, if given, is asked for a name requested but not registered (resolve):
what it returns is built and registered under that name, once however many requests
ask at the same time, and served from then on like any other dataset.
resolve(name)
¶
The pipeline named name: the registered one, else the one the resolver builds
for it (then registered), else None. Exceptions from the resolver propagate.
refresh(name=None)
¶
Rebuild dataset name (every one if None) so its ops' cache_token is read
again, after the state outside their parameters changed. Returns the names whose
digest changed: they get new links, and subscribers and /api/events hear of it.
subscribe(callback)
¶
Register callback(name, pipeline); returns an unsubscribe function.
Ops¶
Bases: BaseModel
A block-wise operation.
Subclass, set name, declare parameters as pydantic fields, and implement apply.
halo: voxels of upstream context needed on every side (int or per-axis tuple). The framework reads the padded region, callsapplyon it, and crops the result.cache: whether this stage's output chunks should be memoized. Turn on for expensive stages (inference) so cheap downstream tweaks (threshold) are free. A pipeline can override it per op:{"op": "gaussian", "sigma": 4, "cache": true}.output_dtype/output_info: describe the result; default is unchanged.output_kind: what the result's values are (ArrayInfo.kind):image,labelormask;None(the default) if the op does not say.cache_token(): what the result depends on beyond the parameters (a file's modification time, weights replaced in place); part of the op's identity.slots: at most this manyapplycalls of this op at once, process-wide (a GPU, a memory-hungry step), queued finest level first;None(the default): no limit.
cached
property
¶
Whether this op's output is memoized: its own setting, else its class's.
input_voxel_size()
¶
The voxel size this op must read, along the data's last axes and in the source's
units: a model trained at one resolution. The pipeline then runs it on one level (the
source's at that size, else one resampled to it), caches what it makes, and makes the
coarser levels by downsampling that. None (the default): it runs on every level.
for_level(info)
¶
This op as it runs on a scale level described by info (the stage's output, on
the input's grid). Ops that work in physical units override it to take the level's
voxel size: a slope in degrees needs the pixel spacing, which doubles from level to
level. The default is the op itself.
apply_at(block, box)
¶
Like apply but told where block sits (box = its halo-padded extent in
voxels of this scale level). Override when the result must depend on position, e.g. to
make per-chunk labels globally unique. Default delegates to apply.
cache_token()
¶
What this op's output depends on outside its parameters, as a string that changes
when that does: a weights file's modification time, a model version. None (the
default): nothing. It is part of digest, so a pipeline built after it changes has
new stage keys and new links, and viewers refetch; DatasetRegistry.refresh
rebuilds a served dataset to read it again.
digest()
¶
Identity of what the op computes: whether it is cached does not change that.
Sources and core types¶
Bases: ABC
Read-only, random-access array. Everything upstream of an op is a Source.
read(box)
abstractmethod
¶
Return data for box (must lie within info.shape), shape == box.shape.
read_padded(box, fill=0, *, edge=False)
¶
Read box even if it pokes outside the array. Out-of-range voxels are fill,
or with edge the nearest voxel inside (so a filter sees no step at the border).
cache_key()
¶
Stable identity used to build cache keys for downstream stages.
Bases: Source
A source materialized chunk-by-chunk via compute_chunk and memoized in an LRU.
read(box) gathers the covering chunks (from cache or freshly computed) and slices.
This is the single mechanism behind both raw-source caching and per-stage caching:
every pipeline stage is exposed to the next one as a ChunkedSource.
chunk(index)
¶
Full (edge-clipped) chunk index as an array of shape info.chunk_box(index).shape.
An ordered list of Source levels, s0 = full resolution. shader is a
Neuroglancer shader the viewer should show the data with (its channel axis then being
a shader channel), for sources whose channels mean something particular.
level_for(voxel_size, rtol=0.01)
¶
The level to read data at voxel_size from (its last axes, in the source's
units), and whether it is at that size: one that is (within rtol), else the
coarsest finer on every axis (resampled down, never up), else level 0.
nearest_level(voxel_size)
¶
The level whose voxel size (its last axes) is nearest voxel_size, by ratio
(the sum over axes of the size's log ratio); the finer of two as near.
Static description of one scale level of a chunked array (C-order axes). kind is
what its values are, for viewers choosing how to show it: image (intensities),
label (segment ids) or mask (inside or not); None if unknown.
chunks_covering(box)
¶
Yield chunk indices whose extent intersects box (assumed already clipped).
rescaled(voxel_size)
¶
This array's extent on voxels of voxel_size along its last axes (as many as
given): the shape rounded up, chunks as they were (in voxels), and the translation
moved so each voxel's position is its centre, as OME-Zarr's is (a voxel twice as big
covers two, its centre half an old voxel on).
Half-open voxel box [start, stop) in C-order axes.
relative_to(other)
¶
Express this box in coordinates whose origin is other.start.
Bases: Generic[T]
Thread-safe LRU keyed by any hashable, bounded by total bytes.
A single shared instance is normally used for every pipeline stage; keys embed the stage hash so that changing a parameter downstream never evicts upstream entries except through normal LRU pressure.
invalidate(prefix=None)
¶
Drop entries; if prefix is given, only tuple keys starting with it.
A request's hold on the work it needs, cancelled when its client stops waiting.
seq is its place in line, taken when the request arrives.
fn(*args) for a request holding claim, in one of slots (run in a server
thread); a request given up on before its turn comes costs nothing.
At most n requests computing at once, each in its own thread, which bounds the
memory and CPU chunk work takes. A request waiting on a queue's work gives its slot up
meanwhile (its thread only waits), so requests that need nothing expensive are never held
up behind ones that do.
Coordinate transformations¶
The model every registration format is read into, the OME-Zarr 0.6 reader, the
deformable solver, and the scene://, warp:// and register:// sources built on them.
Bases: ABC
A function from points in one coordinate system to points in another.
apply(pts)
abstractmethod
¶
Map (n, ndim_in) points to (n, ndim_out) points.
inverse()
¶
The closed-form inverse, or None if there is none.
approximate_inverse()
¶
An inverse that may be estimated numerically; defaults to the exact one.
describe()
abstractmethod
¶
JSON-able identity; equal descriptions mean equal functions (used in cache keys).
A regularly sampled vector field: an array with one vector axis, plus the affine that maps its other (grid) axes to points in the transform's input space.
sample(pts) interpolates the vectors at arbitrary points, reading only the window of
the array those points touch. Outside the field, values are clamped to the nearest edge.
sample(pts)
¶
Vectors at pts (NaN rows, e.g. unconvergent inverse points, stay NaN).
Bases: Transform
y = x + d(x) with d a vector field in the input space (sampled, or
procedural like SwirlField).
Bases: Transform
Numerical inverse of x -> x + d(x): solves x = y - d(x) by fixed-point
iteration, which converges wherever the field's Jacobian has norm below one (true for
the smooth, non-folding warps registration produces). Points that do not converge
(the warp folds there, so no unique inverse exists) come back as NaN, which the
resampler renders as empty rather than as a wrong value.
A procedural displacement field: a twist about an axis, computed exactly from coordinates, so nothing is stored and any parameter can change at any time.
Points at distance r from the axis through centre turn about it by
angle * exp(-(r / radius)**2) degrees: fully on the axis, fading out over about
radius. Only the two plane axes move. Same interface as VectorField.
Several 3-D swirls, their displacements added: swirl k turns points about the
axis axes[k] (a direction, C order) through centres[k] by
angles[k] * exp(-(d / radii[k])**2) degrees, d the distance from its centre.
So each is a ball of twist fading in every direction, and a tilted axis moves points
along all three axes. Where swirls overlap the sum is smooth but no longer a pure
rotation. Same interface as VectorField.
The transformation graph of an OME-Zarr 0.6 scene or single multiscale image.
transform(src, dst, *, approximate=False)
¶
The transform taking points in src to points in dst, composed along the
shortest path of the graph. Edges are walked backwards only when the transform has
a closed-form (or, with approximate, an estimated) inverse.
One transformation object -> Transform. base is the URL of the group whose
metadata holds obj; path parameters resolve against it. ndim_in/ndim_out
are the dimensionalities of the referenced coordinate systems, where known.
fields says how the displacement/coordinate fields it opens are read.
Fit u level by level (fixed[i] against moving[i], coarse to fine).
affine (4x4 or 3x4) maps fixed physical points to moving physical points; the
moving image is sampled at affine(p + u(p)). The control grid spans box (the
fixed image's physical extent, (low, high)); points outside it use the nearest
edge value. Returns snapshots fields evenly spaced over the iterations, from the
starting field to the final one (just the final one when snapshots is 1).
init starts from a field given at the first level's control points, (z, y, x,
3), rather than from zero; ranges normalizes the images with the given
intensity ranges (fixed, moving) rather than the first level's percentiles, so blocks
of one image are all scaled alike; label names the fit in the log.
The fixed-to-moving affine (4x4, physical units) to start from when none is given,
as the browser page's affine.ts finds it. The images' intensity moments are
matched first, centre to centre and principal axis to principal axis, which leaves
each axis's sign open: of the orientations that do not mirror the image, the best
correlated is kept. A mirror image (one axis reversed) is not searched for: on a nearly
symmetric specimen (a brain) it correlates as well as the right rotation, so the images
cannot tell, and a fit cannot reach one from a rotation; give it as an affine. Then the 12 numbers are fitted by gradient ascent on the
normalized cross-correlation, on a copy of at most MAX_FIT_VOXELS voxels after a
coarser one. The moments assume both images show the same whole object. Also returns
the correlation with no affine, after the moments and after the fit.
Stitching¶
Tiles stitched by interest points and RANSAC, and fused, behind stitch:// and the
browser's stitch page.
Bright blobs in block: local maxima of a difference of Gaussians (sigma and
2^(1/4) sigma, BigStitcher's), at least threshold of the intensity range [lo, hi],
to subvoxel precision. Returns (n, 3) voxel positions in block; none within a
voxel of its faces (their maxima are cut off).
Candidate matches (i, j) of points a[i] and b[j]: b[j] is the point
whose descriptors come closest to one of a[i]'s, significance times closer than
any other point's, and the other way round.
The model that most candidate matches p[i] -> q[i] agree with to within
epsilon, from iterations random minimal samples, refitted to its inliers until
they stop changing. Returns the model (None if too few agree) and the inliers.
Every tile's correction (applied after its stage placement) fitted to all the kept
matches at once: link (i, j, p, q) asks tile i's points p to meet tile
j's q (both placed by their stage positions). Each round refits each tile to
where its neighbours put their ends of its matches. Of each group of tiles that links
join, one stays where it is: fixed in its group, the first tile in the others.
Returns the corrections and which tiles were joined to another.
Matches, RANSAC and the global fit from each overlapping pair's interest points
(points[k]: pair k of overlaps, both tiles' points in scene coordinates at
their stage placement). JSON-able: per pair the candidates, inliers and errors; per
tile its correction and its final placement.
Voxels [out_lo, out_hi) of the fused volume on grid: each tile sampled
(trilinearly) where its placement puts it, from blocks[t] (its level's voxels from
a start, or None where it has none of the region), averaged with weights that fall
off near its edges, so the seams do not show.
The tiles of one channel of a BigStitcher project (its XML text; base: the URL
of the folder it is in), each with its stage placement: every view transform but the
outermost ones named drop (the stitching found before, kept as the reference). Reads
projects whose images are OME-Zarr (BigStitcher-Spark's bdv.multimg.zarr loader),
one group per setup and time point.
Tracking¶
An object followed through a time series of labels by its overlap, frame by frame, and
through its divisions into its likely daughters (new objects appearing near where it was: a
guess, not something the labels record), behind the browser's track page and
examples/track_nucleus.py.
The object labelled label at frame t0 followed forward and back (follow),
and forward through its divisions: the daughters of each nucleus that collapses and is
lost are followed too. Returns branches (frame -> record), the first the object's own;
each daughter's first record says mother, the branch it came from.
The object labelled label at frame t0, through frames (counting up or down
from t0); read(t, lo, hi) gives a frame's labels over a box of the level. Yields
(t, record) for t0 and each frame it is found in, until it is lost; a record
says divided when its volume fell to DIVIDED of the frame before's.
The object labelled label in prev, in the next frame nxt (the same box of
the level, from start): the label it overlaps most there, measured, with the overlap
(its share of the object's voxels). None if under MIN_OVERLAP.
Labels of nxt that are new: covered less than NEWBORN by any one label of the
frame before, prev (the same box, from start); within within (physical units)
of centre (level voxels) and of min_volume or more. After a mitosis, which hides
the nucleus for a few frames, these are likely its daughters (a guess: nothing says which
nucleus a new one came from). Nearest first.
Up to two likely daughters of mother (a record, see mother_of): the first new
nuclei (newborns) near it, looked for from frame
lost for gap frames in a box within around it: (frame, record) where each
first appears.
If a branch lost at frame lost had collapsed first (its last volume DIVIDED of
its largest in the dozen frames before, or less), the record at that largest: the
nucleus that went into mitosis. None if it was just lost.
Label label in block (whose voxel 0 is start of the level): its volume
(voxels times their size), centroid and bounding box (level voxels, [lo, hi)), and
whether it touches the block's faces (the block may not hold all of it).
Frontends¶
Bases: ABC
resolve(pipeline, path)
abstractmethod
¶
Map a request path (relative to the dataset+format prefix) to metadata or a chunk.
encode(info, index, block)
abstractmethod
¶
Encode the (edge-clipped) block for chunk index in this format.
resolve_range(pipeline, path, start, stop)
¶
A request for bytes [start, stop) of path, for files read by HTTP Range
requests; None serves the whole file as resolve does.
compute(pipeline, req)
¶
The body for req: its chunk, encoded. Frontends that serve something made from
the data rather than its chunks (meshes) override this.
listing(pipeline, path)
¶
What a directory listing of path shows: the group's metadata keys and one
directory per level, or a level's metadata keys (never its chunks); None if path
is neither, or the format has no directories to list.
relative_scales(pipeline)
staticmethod
¶
Downsampling factor of every level relative to s0, per C-order axis.
pad_to_full(info, index, block, fill=0)
staticmethod
¶
Zarr requires full-size chunks even at the edge; pad with fill.