Model Pipeline
Hugging Face to inference in one flow. Browse, download, quantize, deploy.
Full model lifecycle
Browse HF Hub
Search by family, size, task, quantization type. FTS5 full-text search across 800K+ models.
Download with Progress
Real-time download progress bar. Resumable. Respects rate limits and gated model auth.
MLX Conversion
Convert any Hugging Face model to MLX format via mlx_lm.convert. Automatic architecture detection.
Quantization
2-bit to 8-bit quantization. Pick your quality/speed tradeoff. Mixed precision supported.
One-Click Deploy
Send to any cluster node via asmi /serve/load. No SSH. No config files. One POST request.
Cross-Node Sharing
Transfer models between nodes at 5+ GB/s via mlx_lm.share over Thunderbolt 5 RDMA.
Pluggable storage
Storage backends are auto-discovered at boot. Local SSD, NAS mounts, external Thunderbolt drives — anything the OS can see, the pipeline can use. New volumes appear in the model inventory automatically. No config changes needed.
Same principle applies to compute. When a new node joins the cluster and asmi comes online, its storage and GPU capacity become available to the pipeline instantly. Remove a drive or a node and the inventory updates within one poll cycle.
12 Model API routes
Growing
Models indexed
12
API routes
FTS5
Search engine
4
Storage backends
Auto
Discovery