Back to features

Natural Language Ops

Talk to your cluster. It talks back with rich, real-time data.

Architecture

Cluster tools

Node Health

Check any node's status, CPU, GPU, RAM in real time.

Start / Stop Models

Deploy or undeploy models with a tool-approval gate before execution.

Distributed Inference

Launch JACCL multi-node tensor-parallel serving across the mesh.

RDMA Status

Check InfiniBand link health, ARP peers, and bandwidth across every TB5 cable.

Process Inspector

See running MLX processes, PIDs, RSS memory, and uptime per node.

Model Sharing

Transfer models between nodes at 5+ GB/s over the RDMA fabric.

Running Servers

Find all active inference endpoints across the cluster instantly.

Shell Access

Execute commands on any node. Gated behind tool approval for safety.

System Profile

Chip model, RAM capacity, macOS version, Tailscale status per node.

Web Search

Search the web from inside the chat. Research without switching apps.

Model Info

HuggingFace metadata, community recommended parameters, quant details.

Cluster Scan

Full discovery sweep across all Tailscale peers and TB5 links.

Rich tool previews

Every tool returns structured data rendered by a dedicated React component -- not raw JSON. 21 custom preview renderers turn tool results into something you can actually read: health checks with colored status badges, process lists with CPU and memory usage bars, RDMA topology maps with per-link state indicators, and model deployment progress with live step tracking.

Health badgesProcess tablesRDMA topologyDeploy progressMemory barsModel cardsMetric sparklines+14 more

Auto-Heal

When something breaks in the cluster, the auto-heal pipeline kicks in. The AI detects the failure, diagnoses root cause, proposes a fix, then waits for your approval before applying it. You stay in control -- the AI does the legwork.

Detect

Diagnose

Propose

Approve

Apply

AI proposes: restart mlx_lm.server on m3u1 --model Qwen3.5-27B

ApproveRejectYou decide. The AI waits.

Live observability

The ChatStatusBar sits below every conversation, streaming real-time inference metrics as the model generates. No guessing, no waiting for logs -- you see exactly what the cluster is doing.

TPS42.3tok/s
TTFT184ms
In1,204tok
Out847tok
Reason312tok
Context67%
Step3/5
Voice input available -- press and hold to dictate

Many

Cluster Tools

21

Preview Renderers

Live TPS/TTFT

Observability

Voice + Text

Input

Auto-heal

Recovery