Pricing
Hourly model deployments. Production version should add token pricing, dedicated GPUs and enterprise support.
Llama 3 8B API
Managed vLLM endpoint for chat, coding and general inference.
$0.75/hr
RTX 4090 / A40
Qwen 2 7B API
Multilingual API endpoint for English and Asian-language workloads.
$0.65/hr
RTX 4090
DeepSeek Coder API
Code generation endpoint for developer tools and automation.
$1.90/hr
A100 / H100
ComfyUI Workspace
Browser-based Stable Diffusion and image generation workflows.
$0.85/hr
RTX 4090