Graphics and 3D
AbstractTUI renders real pixels in the terminal: PNG/JPEG images through the best channel the terminal offers, and software-rasterized 3D models (GLB) with textures, lighting, and animation — all with hand-rolled decoders and no GPU requirement. Degradation is always labeled, never silent: when the engine falls back to a lesser channel, the result says so.
Images end-to-end
Bytes to picture
gfx::decode_image(bytes) sniffs the magic bytes (containers lie, bytes
don’t) and decodes PNG or JPEG into a gfx::Bitmap — an owned
RGBA8 image with pixel get/set, nearest and bilinear resize, cropping,
and a box-filter mip chain. Unknown formats reject by name, telling the
caller what does decode (“PNG and JPEG decode, GIF/WebP/AVIF/TIFF do not”
— a message you can show verbatim). Truncated
or hostile bytes produce named errors, never panics; the decoders are
fuzz-hardened.
The JPEG side reads the 8-bit Huffman frames a camera or an image editor actually writes: baseline, extended sequential, and progressive, grayscale or YCbCr, at 4:4:4, 4:2:2, 4:4:0, or 4:2:0, with interleaved or per-component scans and restart markers. Progressive files — what most editors emit by default, and the usual shape of a photo saved for the web — decode through the full spectral-selection and successive-approximation ladder, so the result is the final picture, not a first-pass approximation. Arithmetic-coded, lossless, hierarchical, 12-bit, and CMYK JPEGs reject by name.
Picture to terminal: the capability ladder
The engine picks the best channel the terminal proves it supports —
gfx::choose_channel(&caps.graphics()) (the ladder reads the
graphics view of the capability report) — best first:
| channel | how it draws | moves / resizes | removal | requires |
|---|---|---|---|---|
| kitty graphics | upload once by id, place by escape | cheap re-place, no retransmit | true delete | kitty graphics protocol |
| iTerm2 | full base64-PNG re-emit at the cursor | full re-emit | cells overdraw | OSC 1337 support |
| sixel | paletted raster at the cursor | full re-emit | cells overdraw | sixel + known cell pixel geometry; one shared palette |
| unicode mosaic | colored glyphs (it is cells) | free | free | any terminal |
Capabilities come from detection, not folklore: an instant environment
pass, then an active query probe that can raise and lower the answer.
Run any of the dashboard, viewer3d, or images examples with --caps
to print the report for your terminal.
Three entry points, smallest first
gfx::render_to_cells(bitmap, rect, &caps)— one call. Picks the best mosaic mode for the probed terminal and returns ready-to-blitCellPatches.widgets::Image— the widget:Image::from_path("logo.png")orImage::from_bitmap(Arc<Bitmap>)(theBitmaptype is re-exported beside the widget and in the prelude), withfit(Contain/Cover/Fill/None), alignment, and a mosaic-mode override. The widget always renders mosaic cells: a widget draw closure owns cells, not escape bytes, so pixel-protocol placement lives one level up.gfx::ImageSession— pixel protocols with a lifecycle. Slots are keyed by the caller (SlotKey), content changes are declared by version bump, and each sync emits the minimum traffic the channel allows: kitty transmits once, re-places on move, and deletes on drop; iTerm2 and sixel honestly re-emit their full payload on any change. Bytes flow through anExternalSink(the presenter adapts), and tmux passthrough wrapping is applied automatically when capabilities call for it.SyncOutcometells you whether cells need repainting, bytes were written, or nothing changed.
gfx::present_image / ImageRenderer sit under all three: capability
ladder on top, RenderConfig for the knobs (kitty wire format, placement
z-index, sixel register budget and dithering).
Mosaic modes
Mosaic renders pixels as colored glyphs with a two-colors-per-cell best fit (weighted least squares):
- HalfBlock — 1×2 pixels per cell using
▀. Exact and universal. - Quadrant — 2×2, the 16-glyph quadrant set. Universal glyph coverage.
- Sextant — 2×3, the 64-pattern sextant set. Denser, but its U+1FB00 glyphs need a recent font — explicit opt-in, since no font probe exists and missing glyphs render as tofu.
- Braille — 2×4, dots by luminance threshold. Structure rather than color; the strongest choice on monochrome-class terminals.
MosaicMode::auto(&caps) picks for you and returns the reason as a label:
non-UTF-8 locales get HalfBlock (U+2580 survives most legacy codepages),
monochrome terminals get Braille, color terminals get Quadrant.
Animated pictures (and why video is not one of them)
The engine plays a frame sequence in the cell grid. Two formats decode in-tree: animated GIF and APNG. Both are permissively licensed, free of patent pools, need no dependency the crate does not already have, and are small enough to review.
#![allow(unused)]
fn main() {
use abstracttui::prelude::*;
use abstracttui::widgets::{AnimatedImage, ImageFit};
fn build(cx: Scope) -> View {
AnimatedImage::from_path("loading.gif")
.fit(ImageFit::Contain)
.view(cx)
}
}
gfx::decode_animation(bytes) is the decoder behind it, magic-routed
like decode_image, and a still decodes to a one-frame animation — so
one code path shows any picture, moving or not. Animation::frame_at
answers “which frame shows now” for a caller driving its own clock.
Video does not decode here. .mp4, .mov, .avi, .webm and
.mpg carry H.264, H.265, VP9, or MPEG-4. Those are patent-pooled, and
a correct real-time decoder for any one of them is larger than this
entire crate, so the engine does not pretend: the container is
recognized, named, and refused with the line that converts it into
something that does play.
mp4/mov: video is not decoded (animated GIF and APNG are).
Convert with: ffmpeg -i IN -vf 'fps=12,scale=480:-1' OUT.gif
To play video as it is, decode it OUTSIDE the engine and feed frames
in — the same shape a terminal app uses for audio. Spawn
ffmpeg -i clip.mp4 -f rawvideo -pix_fmt rgba -, read frames on a
thread into a reactive::bounded_source, build a gfx::Bitmap per
frame, and show each one with Image::from_bitmap. Kill the child in
the scope’s on_cleanup: an orphaned decoder holds the file and the
CPU. That dependency is the CALLER’s, deliberately — the engine ships no
bindings to ffmpeg, claims no video support, and works without it.
What playback costs. A moving picture is the one thing here that cannot be idle, so the cost is stated rather than hidden: each frame arms ONE timer for that frame’s own delay, so a playing clip costs one wakeup per frame, nothing between frames, and nothing at all once it is paused or finished. Bytes are the channel’s business — a mosaic repaint is a cell diff (only what changed), which makes the universal path the cheapest one for motion, while kitty, iTerm2, and sixel each re-send a full payload per frame.
Getting more resolution out of a picture
A mosaic cell carries one glyph and two colors, so a picture’s resolution is the glyph family’s subpixel density — and the family is the first knob to reach for.
widgets::Image, markdown image blocks, and gfx::render_to_cells all
follow MosaicMode::auto, so a UTF-8 color terminal already draws at
quadrant density (2×2 per cell). Pinning Sextant buys another step
where you know the font carries the Unicode 13 block sextants:
#![allow(unused)]
fn main() {
use abstracttui::gfx::MosaicMode;
use abstracttui::widgets::{Image, ImageFit};
Image::from_bitmap(photo)
.fit(ImageFit::Contain)
.mode(MosaicMode::Sextant) // 2x3 subpixels — check your font first
.view(cx)
}
Measured on a 1200×1599 photograph drawn into a 24×32 cell pane, scoring each family against one common reference:
| family | subpixels per cell | PSNR |
|---|---|---|
| HalfBlock | 1×2 | 20.5 dB |
| Quadrant | 2×2 | 23.5 dB |
| Sextant | 2×3 | 24.3 dB |
Two more levers, in order of effect:
- Give the image more cells. Resolution is cells × subpixels, so a pane twice as wide is worth more than any family change: the same photo at 40×50 cells gains ~1.4 dB per family over 24×32.
- Reach the pixel protocols. kitty, iTerm2, and sixel send real
pixels, which no mosaic family can match. That path is
gfx::ImageSessionat the app level (a widget draw closure owns cells, not escape bytes) — see the three entry points above.
Optional Floyd–Steinberg dithering (serpentine error diffusion) can pre-quantize the source to a palette before cell fitting — worth it when the output terminal is 256- or 16-color, where straight quantization would band gradients. Sixel emission has its own configurable dithering.
Images under tmux
Inside tmux, graphics protocols are off by default because tmux swallows
them unless the user set allow-passthrough on — a setting invisible from
the environment. The engine verifies passthrough per session with a
wrapped round-trip probe and only then enables the kitty/iTerm2 paths,
wrapping every payload automatically. Known cosmetic limit: tmux cannot
reflow passthrough images across scrolling or pane splits. Mosaic works
everywhere regardless.
Verifying image support on your terminal
Two commands answer “what does my terminal actually support?”:
cargo run --example caps # the capability report
cargo run --example images # see it: mosaic families + protocol placement
caps renders the live capability set — the probe’s upgrades appear on
screen moments after launch, and the images via line names the channel
the ladder picked (kitty graphics protocol, iterm2 inline images,
sixel, or unicode mosaic). Apps can read the same facts through
use_caps(cx) / current_caps(). images then shows the result: the
four mosaic glyph families side by side, and p places the same picture
through the chosen pixel protocol, labeled with the channel in the
footer.
What to expect per terminal: kitty, WezTerm, and Ghostty take the kitty
graphics path; iTerm2 and VS Code’s terminal take OSC 1337 inline
images; foot and mlterm take sixel; Terminal.app and most others render
unicode mosaic — which is not a failure but the universal fallback, at
character-cell resolution. Under tmux, protocols engage only when the
passthrough probe proves allow-passthrough on (see above). If a
protocol row reads yes but you see mosaic, check cell pixel size —
without pixel geometry the ladder stays conservative.
Not an image: hand-drawn vector strokes
Charts, gauges, and hand-rolled traces do not go through the image
pipeline at all. The public sub-cell canvas (DotCanvas, in the
prelude) draws braille/quadrant dot grids with line/bezier/arc strokes
and eighth-block fills — the same layer the shipped charts render
through, and what the diagram extensions stroke their edges with.
The full surface (dot-space model, primitives, the cell-color rule) is
api.md § canvas; for
node-and-edge diagrams, prefer the abstracttui-graph /
abstracttui-mermaid crates over hand-stroking
(graphs-and-diagrams.md).
Not an image: text drawn several cells tall
gfx::bigtext rides the same mosaic encoder described above, but its
source is the engine’s embedded 8x16 font rather than a decoded picture.
It exists because a terminal has one font size: a bigger glyph means
spending more cells and subdividing them, so a banner heading or an
oversized icon is a rasterization problem, not a typography one.
Because it reuses the encoder, the capability ladder and every mosaic mode apply unchanged — but the best mode differs from the image case. Braille packs the most subpixels per cell and terminals draw its dots with visible gaps, which reads well for photographs and poorly for letterforms; sextants are solid ink at 2x3 and usually read better as type.
Two rules the surface makes explicit, both easy to get wrong by
guessing: how small you can go is a MEASUREMENT and not a constant —
bigtext::smallest_clear(mode, content) takes it, because uppercase,
mixed text and icons bottom out at different sizes and the mosaic mode
moves all three — and a square icon needs twice as many columns as rows
because a cell is about 1:2, which GlyphScale::square(rows) does for
you. The sharp edge worth knowing before you pick: GlyphScale::FLOOR
(4x3) is clear for everything in braille and only MARGINAL for mixed
text in sextant, the mode this page just told you to prefer for type;
smallest_clear sends you to 4x4 there.
The second edge is the one that produced a bug report, and it took two
fixes. Pairwise distance measures whether two rasterizations DIFFER, and
more columns always buys more difference — so a cheapest-clear search
with no bound walks to the widest scale it can and calls it best. It used
to return 6x2 for braille text, where a filled ● reduces to six
near-solid cells and looks like a white bar; the measure was right that
one bar differs from the next by 16 subpixels and had no way to notice
that neither was its glyph any more.
The search is now bounded to twice the glyph’s natural aspect in either
direction (GlyphScale::within_aspect_band), which also lets it reach
past six columns: MosaicMode::HalfBlock clears at 7x5 for mixed text,
and this page used to say it never cleared at all. And legibility
now measures the shape as well as the distance —
bigtext::fidelity_loss compares what is drawn against the same glyph
at its natural proportions, so a scale you name yourself is graded
honestly too: 6x2 comes back Legibility::Distorted, not Clear. The
full surface, its
limits (no accented glyphs) and the measurements behind the floor are in
api.md § gfx::bigtext;
cargo run --example bigtext shows every scale and symbol set on your
own terminal.
3D end-to-end
The five-line hello
#![allow(unused)]
fn main() {
use abstracttui::three::{self, Framebuffer, SceneRenderer};
let view = three::quick_view("model.glb")?; // load + framed camera + light
let mut fb = Framebuffer::new(160, 96);
SceneRenderer::new().render(&view.scene(), &mut fb);
// fb -> mosaic cells via MosaicRenderer, or use Viewport3D below.
}
quick_view (and quick_view_bytes for in-memory GLB data) returns a
QuickView with public model, camera, light, and stats fields —
adjust the camera and light freely between frames, then call .scene().
stats reports decode cost (texture decode dominates on textured models;
a 2048² JPEG-textured asset loads in the ~100 ms class — show a loading
state around it). look_from(yaw, pitch) re-frames the camera on the
model’s bounds: the “reset camera” a viewer needs.
What loads (the GLB subset)
Binary GLB containers with embedded buffers; TRIANGLES primitives;
positions, normals, UVs, and vertex colors; u8/u16/u32 indices or
non-indexed geometry; node TRS and matrix hierarchies; multiple scenes
(the default scene wins); baseColorFactor and baseColorTexture
(embedded PNG or JPEG); emissiveFactor; smooth-normal generation on
request; and a 2-million-triangle budget enforced from metadata before
decode.
Rejected by name: sparse accessors, Draco/meshopt compression, non-triangle primitive modes, and out-of-range anything. Labeled degradations (the model loads, with a warning): external URIs, unsupported texture maps (normal/metallic-roughness/occlusion), morph weights, and CUBICSPLINE animation channels (skipped).
Scene, camera, light
Scene<'_> borrows a Model and carries a Camera (orbit-style: target,
yaw, pitch, distance, vertical FOV, near/far), a Light (directional:
direction vector or spherical from_angles, ambient + diffuse terms), a
background color, and a double_sided flag.
Culling defaults differ by entry, deliberately: bare Scene::new culls
back faces (procedural meshes are consistently wound), while Viewport3D
and QuickView::scene() render double-sided (real-world GLB exports are
not, and holes read as bugs). Flip double_sided explicitly when the
other trade-off fits.
The Viewport3D widget
#![allow(unused)]
fn main() {
let vp = Viewport3D::new(Arc::new(model))
.orbit(yaw, pitch, zoom) // plain floats each build; signals live app-side
.mode(MosaicMode::HalfBlock)
.animate(0, t) // play clip 0 at time t (loops; static = rest pose)
.light_angles(azimuth, elevation)
.fog(0.15)
.on_orbit(move |dyaw, dpitch| { /* write yaw/pitch signals */ })
.on_zoom(move |steps| { /* write zoom signal */ })
.element(&tokens);
}
The widget is pure over its props: same props, same pixels. Left-drag
orbits (the pointer is captured for the drag, so fast drags keep steering
outside the rect), the wheel zooms — but the widget only reports deltas
through on_orbit/on_zoom; the app owns camera state and clamping.
element(&TokenSet) takes no scope because the widget holds no reactive
state. The default layout grows into whatever region the parent hands it
(a viewport has no intrinsic size to measure), so an un-layout()ed
viewport is visible by construction; pass .layout(...) to size it
explicitly. Buffers persist inside the draw closure, so a steady-state
repaint allocates nothing. light, background, spin (caller-driven
auto-rotation), and cull_backfaces round out the builder.
For a complete interactive viewer, run
cargo run --example viewer3d -- model.glb.
Animation playback
Model::animations() lists the clips; sample_pose_full(clip, t, &mut Pose) produces per-instance world matrices plus per-skin joint matrices —
pure in t, clamped to the clip’s keyframe range (loop with
t % clip.duration()), and allocation-free in steady state (the Pose
scratch is reused across frames).
Supported interpolation: LINEAR and STEP; rotations use
shortest-path nlerp. Skinning: up to 4 joints per vertex
(JOINTS_0/WEIGHTS_0), linear blend, sanitized at load — out-of-range
weighted joints reject, drifted weight sums renormalize with a label. An
animated, skinned test asset ships in the repository
(src/three/fixtures/animated_bar.glb).
Textures and mip-mapping
Base-color textures decode through the same image pipeline (embedded PNG / JPEG) and build a box-filter mip chain. The rasterizer picks a mip level per triangle from the texels-per-pixel ratio, with bilinear sampling within the level; wrap mode is REPEAT.
The boot splash
An optional two-second identity animation for app startup, played before
your first frame. boot::should_splash(&caps) is the production gate: it
returns the reason to skip when the render handle is not a tty, when
ABSTRACTTUI_NO_SPLASH is set (any value except 0, so wrapper scripts
can force-enable), when NO_COLOR is set, when TERM=dumb, or when the
capability report itself says the terminal is dumb. Respect the reason —
it is ready-made for a log line.
The sequence runs 2.0 s in four beats: arrival (three planes fly in, staggered, on an ease-out curve), alignment (at 0.9 s the planes lock into the mark and a 12-spark burst fires), reveal (at 1.4 s the wordmark tracks open from 4 cells of letter-spacing to 1), hold (settle, then done). Any key skips with a fast 120 ms fade, and a hard 2.5 s wall cutoff bounds the whole thing.
Two render paths read the same identity constants: a 3D path (the mark
rendered by the three rasterizer, chosen on truecolor terminals) and a
pure-cell 2D path with its own particle field (everywhere else). Try both:
cargo run --example splash (--3d / --2d to force one).
Honest limits
- JPEG: 8-bit Huffman frames only — baseline, extended sequential, and progressive. Arithmetic coding, lossless, hierarchical, 12-bit precision, and CMYK reject by name. Chroma upsampling is nearest-neighbour, which costs a code or two against a smooth upsampler at terminal resolutions. Scan component selectors are validated against the frame header; malformed scans reject rather than decode wrong.
- PNG: 8-bit depths, no interlacing (Adam7 rejects by name).
- Animation: animated GIF and APNG decode in-tree. Video does
not —
.mp4,.mov,.avi,.webm,.mpgare recognized, named, and refused with a conversion command, because their codecs are patent-pooled and far larger than this crate. Frames are held decoded in memory (budget: 64 Mpx across a sequence), so a long clip belongs on the external-decoder path rather than in aVec. - Sixel: one palette per emission — multiple live sixel images recolor each other. Prefer one sixel image per screen.
- iTerm2/sixel have no placement model: any move or resize re-emits the full payload; only kitty gets placement escapes and true deletes.
- Pixel protocols are verified byte-for-byte against the protocol specifications and a protocol state model, not against every live terminal emulator; mosaic is the universal, always-correct path.
- Screenshots: cells under a kitty/iTerm2/sixel placement are not the picture — captures export those regions as labeled veils rather than fake cells, while mosaic images capture as themselves (they ARE cells). See api.md § “Screenshots & captures”.
- Animation: LINEAR/STEP only; CUBICSPLINE channels and morph weights skip with labels; rotation interpolation is nlerp, not slerp.
- Skinning:
JOINTS_0/WEIGHTS_0only (4 joints per vertex), linear blend, no inverse-transpose normal handling (an approximation under non-uniform scale). - Textures: base color only; other maps are labeled and ignored; wrap is REPEAT (per-sampler modes are not read). Mip LOD is per-triangle, not per-pixel.
- Mosaic: two colors per cell, by construction; braille carries structure, not color.
- Rasterizer: near-plane and guard-band clipping, top-left fill rule, perspective-correct depth and UVs; vertex-color interpolation is screen-linear (invisible at cell scale).
- Performance numbers are load-sensitive: the envelope below is from an idle machine; medians inflate several-fold under host contention.
Performance envelope
Measured medians, release build, on a quiet machine (ms/frame):
| asset | triangles | 160×96 | 320×192 |
|---|---|---|---|
| synthetic sphere (untextured, gouraud) | 16,128 | 0.76 | 0.97 |
| helmet (JPEG textured + mips) | 15,452 | 1.31 | 1.82 |
| helmet (untextured) | 15,452 | 1.18 | 1.59 |
| x-wing (PNG textured + mips) | 119,999 | 7.53 | 8.21 |
| skinned sphere (animated, all vertices blended) | 65,024 | 2.90 | 3.26 |
The renderer is vertex-bound at cell scale: 4× the pixels costs +9–39%, while 7.4× the triangles costs ~5.7×. Rule of thumb: assets up to ~20k triangles render in well under 2 ms anywhere; a 120k-triangle asset fits a 30 fps budget with 3–4× headroom on one core.
Mosaic conversion adds, for a 200×60-cell target (worst case): half-block ~50 µs, braille ~0.5 ms, quadrant ~1 ms, sextant ~3.7 ms.
Reproduce on your machine:
cargo test --release -- --ignored perf_three_envelope --nocapture
cargo test --release -- --ignored perf_mosaic_200x60 --nocapture