Core concepts
The gateway maps framework capabilities to REST resources under /api.
Resources
| Resource | Purpose |
|---|---|
| Account | Customer profile, permitted providers, available models, and usage |
| Agents | Named LLM configurations with instructions and optional capabilities (web search, RAG, custom tools) |
| Chat | Run a single agent in a one-shot or multi-turn conversation; continue paused tool runs |
| Conversations | List, retrieve, and delete persisted conversation histories |
| Model presets | Reusable model configurations (provider, deployment, temperature, token limits) |
| Memory stores | Customer-scoped long-term memory buckets attached to agents |
| Vector databases | Firestore-backed embedding stores for retrieval-augmented generation |
Typical workflow
- Create a model preset — configure provider, deployment, and parameters.
- Create an agent — attach instructions, a model preset, and optional capabilities.
- Chat — send messages via
POST /api/chatfor one-shot or multi-turn runs. - Persist conversations — load and manage histories via the Conversations API.
Multi-tenancy
Each API key maps to a customer_id. Data is isolated per customer. Some customers use infrastructure overrides (custom GCP project, Firestore database) resolved automatically per request.