Pricing

Hourly model deployments. Production version should add token pricing, dedicated GPUs and enterprise support.

Llama 3 8B API

Managed vLLM endpoint for chat, coding and general inference.

$0.75/hr

RTX 4090 / A40

Qwen 2 7B API

Multilingual API endpoint for English and Asian-language workloads.

$0.65/hr

RTX 4090

DeepSeek Coder API

Code generation endpoint for developer tools and automation.

$1.90/hr

A100 / H100

ComfyUI Workspace

Browser-based Stable Diffusion and image generation workflows.

$0.85/hr

RTX 4090

Pricing | Sanonix