Simple pay-as-you-go pricing for the Assistant Cortex cloud
Your Assistant Cortex instance comes with a marketplace of prebuilt applications and AI tools, and one simple billing model: add funds to your balance, and monthly seat fees and per-use charges are deducted as you go. No contracts, no per-feature tiers, and every new instance starts with a $20 credit.
Need more than an instance? The same codebase that runs our cloud can be licensed to build your own customized SaaS product — see Enterprise Deployment below.
Pay As You Go
Install the applications and AI tools you want from the marketplace, then pay only for what you use. Add funds to your account and monthly and per-use charges are deducted from this balance. Each Assistant Cortex instance comes with a $20 credit to try the product. Contact sales to get your instances today.
CC Payments
We accept all major credit cards including American Express, Visa, Mastercard and Discover. $5 minimum when adding funds to account.
Auto Refill
You can configure your Assistant Cortex instance to automatically refill its balance when a configurable threshold is met.
Monthly Fees
Some services provided by Assistant Cortex are billed on a subscription basis. These services are billed per seat.
User Accounts
Each user account has a $39.99 monthly fee; this fee is deducted each month on the day the user account was created. The first admin account is free until your initial $20 credit is used up — after that it bills as a standard seat.
Phone Numbers
Each phone number has a $3.99 monthly fee; this fee is deducted each month on the day the number was allocated.
Per Use Pricing
Some services offered by Assistant Cortex are billed on a per usage basis. These include AI models and communication services like text messaging.
Assistant Cortex Model Pricing
View pricing information for all available AI models and APIs.
| Model | Category | Input Price | Output Price |
|---|---|---|---|
| Ace Step | Other | 0.000200 | 0.000200 |
| Claude Sonnet 4 | Language Models | 3.60 | 18.00 |
| Flux Fill Pro | Image Generation | — | 0.0200 |
| Flux Kontext Dev | Image Generation | 0.0480 | 0.0480 |
| Flux Kontext Pro | Image Generation | — | 0.0200 |
| Flux Pro 2 Edit | Image Generation | 0.0480 | 0.0480 |
| Flux Pro | Image Generation | — | 0.0200 |
| Gemini 3.1 Pro Preview | Language Models | 2.00 | 12.00 |
| Gemini 3.7 Flash | Language Models | 0.7500 | 3.75 |
| Gemini Flash | Language Models | 0.1200 | 0.4800 |
| Gemini Pro | Language Models | 0.1200 | 3.60 |
| Hunyuan 3D v3.1 Pro – FAL.ai | 3D Generation | — | 0.3750 |
| MiniMax H3 Max Reference-to-Video | Other | — | 0.0800 |
| Mistral Large | Language Models | 0.5000 | 1.50 |
| OpenAI GPT-4o | Language Models | 0.000600 | 0.000600 |
| OpenAI GPT-5 | Language Models | 0.7500 | 6.00 |
| OpenAI GPT-5.4 Mini | Language Models | 0.3600 | 0.3600 |
| OpenAI GPT-5.5 | Language Models | 5.00 | 30.00 |
| OpenAI GPT-5.6 Sol | Language Models | 4.00 | 20.00 |
| OpenAI GPT-6 Astra | Language Models | 10.00 | 50.00 |
| Claude Opus 4.8 | Language Models | 5.00 | 25.00 |
| Claude Opus 5 | Language Models | 5.00 | 25.00 |
| Claude Sonnet 4 | Language Models | 0.001200 | 0.003600 |
| DeepSeek 31 | Language Models | 0.3400 | 0.5000 |
| DeepSeek V4 Pro | Language Models | 0.6600 | 1.98 |
| GPT-6 Astra | Language Models | 10.00 | 50.00 |
| Qwen 3.5 | Language Models | 0.3900 | 2.34 |
| Qwen 3.8 | Language Models | 2.00 | 6.00 |
| Qwen 3.8 Max | Language Models | 2.00 | 6.00 |
| Qwen3 235B | Language Models | 0.2800 | 2.34 |
| Qwen3 Coder | Language Models | 0.2800 | 2.34 |
| Qwen Image Edit 25511 | Language Models | 90000.00 | 90000.00 |
| Hyper3D Rodin v2.5 – FAL.ai | Other | — | 0.4000 |
| Tripo 3D | 3D Generation | — | 0.4000 |
| VibeVoice 7B – FAL.ai | Audio Models | — | 0.000667 |
| VibeVoice TTS | Audio Models | — | 0.0480 |
| Wan 2.1 1B | Other | 0.2400 | 0.2400 |
| AceStep V1.5 Turbo | Audio Models | 0.000240 | 0.000240 |
| CLIP ViT Large Patch14 | Vision Models | 0.000264 | 0.000264 |
| Dream Shaper Image Gen | Image Generation | 0.002900 | 0.002900 |
| Flux 1 Dev Kontext | Image Generation | — | 0.0200 |
| Flux 2 Dev Klein 4B | Image Generation | 0.009000 | 0.009000 |
| Flux 2 Dev Klein 9B | Image Generation | 0.009000 | 0.009000 |
| GLM 4.5 Air | Language Models | 0.1560 | 1.02 |
| GLM 4.5 Air Bartowski | Language Models | 0.7200 | 2.64 |
| Groq GPT-OSS 120B | Language Models | 0.1500 | 0.6000 |
| Groq Qwen 3.8 27B | Language Models | 0.8000 | 4.00 |
| Hunyuan 3D 2 | 3D Generation | 0.1920 | 0.1920 |
| Intel DPT Large | Other | 0.000264 | 0.000264 |
| Juggernaut XL Image Gen | Image Generation | — | 0.0200 |
| Laion CLAP Larger CLAP General | Other | 0.000264 | 0.000264 |
| LTX-2 | Image Generation | — | 0.2000 |
| Media Utils | Other | 0.003000 | 0.003000 |
| Meta SAM2.1 Hiera Large | Vision Models | 0.0252 | 0.0252 |
| OWLViT Base Patch16 | Vision Models | 0.000264 | 0.000264 |
| OWLViT Large Patch14 | Vision Models | 0.000264 | 0.000264 |
| Qwen3 30B A3B | Language Models | 0.000400 | 0.000200 |
| Qwen3 32B | Language Models | 0.001200 | 0.000800 |
| Qwen3 Coder Next 80B | Language Models | 0.2800 | 2.34 |
| Qwen3 Embedding 0.6B | Language Models | — | 2000.00 |
| Qwen 2.5 VL 32B | Language Models | 0.001200 | 0.000800 |
| Qwen 2.5 VL 7B | Language Models | 0.2800 | 2.34 |
| Qwen 3 VL 32B | Language Models | 0.2800 | 2.34 |
| Qwen 3 VL 4B | Language Models | 0.3120 | 3.12 |
| Qwen 3 VL 8B | Language Models | 0.2800 | 2.34 |
| Qwen Image Edit | Language Models | 90000.00 | 90000.00 |
| Qwen Image Edit 2509 | Language Models | 90000.00 | 90000.00 |
| ResNet-50 | Vision Models | 0.000264 | 0.000264 |
| SAM Audio Large | Vision Models | — | 0.0200 |
| Stable Audio Open 1 | Image Generation | — | 0.0200 |
| Tesseract OCR | Vision Models | 0.000264 | 0.000264 |
| VibeVoice ASR Base | Audio Models | — | 0.0000070000 |
| vLLM GLM Air 35 | Language Models | 0.7200 | 2.64 |
| vllm_qwen3_5_35b_a3b_1 | Language Models | 0.2500 | 1.00 |
| vLLM Qwen Qwen3.5 35B A3B | Language Models | 0.4000 | 4.00 |
| Wan 2.2 I2V | Other | 0.1920 | 0.1920 |
| Whisper Medium | Audio Models | — | 0.0000070000 |
| Whisper Medium EN | Audio Models | — | 0.0000070000 |
| Whisper Small EN | Audio Models | — | 0.0000070000 |
| Whisper Tiny EN | Audio Models | — | 0.0000070000 |
| YOLOS Small | Vision Models | 0.000264 | 0.000264 |
| ZImage Turbo | Image Generation | 0.006000 | 0.006000 |
Storage
Your instance’s stored data is metered hourly and billed from your balance as a simple daily charge, prorated per gigabyte-month — you only pay for what you actually store.
| File storage (workspaces, uploads, generated media) | $0.023 per GB-month |
| Database storage (MySQL) | $0.115 per GB-month |
| Search & embeddings storage (Elasticsearch) | $0.122 per GB-month |
Development Services
We can build any custom application, AI ability or agent you need. Our team of veteran developers are here to help make the most out of Assistant Cortex.
Coding Sandbox
Assistant Cortex was built from the ground up to be a self-coding application that can generate advanced applications and AI abilities on its own. With our Coding Sandbox application, you can provide a detailed prompt, and our auto coder will do its best to create the application or AI ability.
Human Coders
For advanced business logic and the hardest problems, a human coder is still required. Our team is AI first and use our internal stack to build what you need as quickly as possible. Contact us for a free quote on the work — finishing a Coding Sandbox application a client has started usually takes only a few hours.
API Providers
Assistant Cortex supports an ever-growing number of cloud providers. These range from large language models to all-encompassing APIs like those provided by HuggingFace, plus multi-provider routers and your own custom-hosted endpoints.
HuggingFace API
Access the HF serverless inference APIs.
Mistral API
Access Mistral’s Large, Medium and Small LLMs.
Google API
Access Google’s Gemini Pro and Flash LLMs.
Claude API
Access Anthropic Sonnet, Opus, and Haiku LLMs.
OpenAI API
Access OpenAI’s frontier LLMs as well as their text to speech, automatic speech recognition and text embeddings models.
OpenRouter API
Access Claude, DeepSeek, Qwen and other LLMs through OpenRouter’s unified API.
AWS Bedrock
Access Claude, Llama, Nova, Mistral and other foundation models via AWS Bedrock.
AWS SageMaker
Connect your own custom-deployed LLM, image or speech endpoints on AWS SageMaker.
Fal.ai
Access Flux, Qwen and Wan models on Fal.ai for image, audio and video generation.
Bring Your Own Key
Already have accounts with the providers above? You can connect your own API keys to your Assistant Cortex instance and run inference through them, paying the provider directly at their rates instead of our per-use model pricing. It’s the easiest way to save on inference costs for high-volume workloads. Contact sales for details.
Your Keys, Their Rates
Add your own keys for providers like OpenAI, Anthropic, Google, Mistral, OpenRouter and Fal.ai. Calls made with your keys are billed by the provider directly, so you pay exactly what they charge.
Mix and Match
Bring a key for one provider and stay pay-as-you-go for everything else. The rest of your billing — seats, phone numbers, storage and locally hosted models — works exactly the same.
Enterprise Deployment
The Assistant Cortex codebase can be licensed to build fully customized SaaS products under your own brand. You get the complete platform — user management, billing, the AI framework, the application marketplace and the coding tools — as your foundation, so your team ships a polished product in weeks instead of building the hardest parts from scratch. We tailor the licensing, hosting and development support to your project; contact sales to talk through what you want to build.
Both of the products below are live customized SaaS builds running on the Assistant Cortex platform.
Ready to get started?
Every Assistant Cortex instance comes with a $20 credit to try the product.

