Real-time analytical
query on operational data

Deploy analytics replicas alongside your operational databases to give apps and agents fast, sandboxed access to data. Open source and deployable anywhere.

Spice AI analytics replica architecture for real-time operational data

Make your data AI-ready

Sub-second query performance across sources, up to 80% lower lakehouse spend, and increased data resiliency.

Lower is better

0x

up to 100x faster queries

Lower is better

0%

up to 80% lower data lakehouse spend

Higher is better

0x

2x increase in data reliability for critical workloads

Powered by

Fast, physically isolated access to enterprise data

A single SQL endpoint for retrieving, accelerating, searching, and reasoning over operational and analytical data. Incrementally adoptable and built on object storage.

Analytics Replica
SQL Federation & Acceleration
Hybrid Search
Embedded AI Inference
Analytics Replica

Real-time analytics without impacting production

Deploy an analytics replica in minutes for fast, isolated queries on real-time operational data.

Analytics on operational data

Spice AI analytics replica architecture overview
SQL Federation & Acceleration

Fast and federated access to all of your data

Connect to and query operational databases, data lakes, and warehouses across the enterprise. Materialize and accelerate working sets in-memory or on disk for millisecond access.

SQL federation and acceleration

Spice AI SQL federation and data acceleration architecture
Hybrid Search

Combine keyword, vector, and full-text search in SQL

Run hybrid search with standard SQL. Spice combines structured filters, vector similarity, and keyword matches in one query, then ranks the merged results by relevance for search-driven apps.

Hybrid SQL search

Spice AI hybrid SQL search combining keyword, vector, and full-text search
Embedded AI Inference

Call LLMs directly from the query layer

Call hosted or local LLMs inline using SQL UDFs or natural language. Translate text, generate summaries, classify entities, and augment query results on your enterprise data without leaving the Spice runtime.

LLM inference in Spice

Spice AI embedded AI inference with LLM calls from SQL

Proven in production

From messaging platforms to security systems, companies like Twilio and Barracuda rely on Spice to deliver low-latency apps & AI agents at scale.

Twilio logo
Barracuda Networks logo
NRC Health logo
Summation logo
Basis Set Ventures logo
Peter Janovsky

“Spice opened the door to take these critical control-plane datasets and move them next to our services in the runtime path.”

Peter Janovsky

Software Architect, Twilio

Darin Douglass

0x

Faster queries

“It just spins up and works, which is really nice. The responsiveness is amazing, which is a huge gain for the customer.”

Tim Ottersburg

“Partnering with Spice AI has transformed how NRC Health delivers AI-driven insights. By unifying siloed data across systems, we accelerated AI feature development, reducing time-to-market from months to weeks - and sometimes days. With predictable costs and faster innovation, Spice isn't just solving some of our data and AI challenges - it's helping us redefine personalized healthcare.”

Tim Ottersburg

VP of Technology, NRC Health

Ramachandra Ramarathinam

“Spice is the data plane behind Summation. Every connector we ship, from Snowflake to a generic REST-as-a-table, collapses into one SQL surface, which is what makes our AI agents portable across a customer's stack.”

Ramachandra Ramarathinam

CTO, Summation

Rachel Wong

“Spice AI grounds AI in our actual data, using SQL queries across many data sources. This brings accuracy to probabilistic AI systems, which are very prone to hallucinations.”

Rachel Wong

CTO, Basis Set

The data plane for enterprise AI

Spice combines SQL query, vector and keyword search, and inline LLM calls in one runtime. AI agents use all three through a single interface.

Spice Cloud dashboard showing a query result

Control and optionality at every layer

Run real-time apps and AI agents on existing data, with control over where Spice runs, what each agent can read, and how every request performs.

Deploy anywhere icon

Deploy
anywhere

Run Spice.ai Open Source locally, at the edge, or on the fully managed Spice.ai Cloud Platform. Lightweight, portable, and designed for scale.

Edge-to-cloud deployments

AI sandboxing and security icon

AI sandboxing
& security

Provision isolated, least-privilege datasets for apps and agents with zero direct database access. Keep governance intact while enabling RAG, agents, and AI workflows.

Secure AI sandboxing

Distributed observability icon

Distributed
observability

Perform end-to-end tracing across SQL, embeddings, search, and LLM calls. Debug, measure latency, and prove ROI from a single view.

Observability docs

Watch Spice in action

Demos from the founders and engineers who built it.

FAQs

Common questions about what Spice AI is, how it works, and how teams run it.

Spice AI builds Spice, an open-source engine for SQL query, search, and LLM inference over enterprise data. Applications and AI agents use one SQL endpoint to query databases, data lakes, and warehouses. Spice can also keep an analytics replica of an operational database, so analytical reads do not load production. Teams run Spice on their own infrastructure or use the managed Spice Cloud Platform.

Spice is written in Rust and uses Apache DataFusion as its SQL query engine. Data moves through the runtime in the Apache Arrow columnar format. Apache Ballista runs distributed queries across a cluster. The Spice Cayenne accelerator stores local copies in the Vortex columnar format.

No. Spice runs next to existing databases, data lakes, and warehouses, and those systems stay the source of truth. Applications keep writing to the operational database. Spice serves reads through SQL federation and acceleration, from the source or from a local accelerated copy. Teams can add Spice one dataset at a time, without a data migration.

AI agents query Spice through three interfaces: a SQL API, an OpenAI-compatible chat API, and an MCP server. The SQL API runs queries that the agent writes. The chat API lets a model call built-in SQL and search tools. The MCP server connects agent frameworks that use the Model Context Protocol.

Spice runs on a laptop, on-premises, at the edge, or in any cloud. Teams deploy it as a sidecar next to an application, as a shared microservice, or as a distributed cluster on Kubernetes. The Spice Cloud Platform runs Spice as a managed service. See edge-to-cloud deployments for each option.

Yes. Spice.ai Open Source is free under the Apache 2.0 license, with community support on GitHub and Slack. Spice Cloud adds a managed service with subscription and usage-based plans. Enterprise plans add commercial licensing, 24/7 support, and SLAs. See pricing to compare the plans.

Twilio, Barracuda Networks, NRC Health, Basis Set, and Summation run Spice in production. The Spice 1.0 announcement from January 2025 reports P99 query times under 5 milliseconds for the Twilio messaging control plane. The Barracuda case study, published in June 2026, reports data lake query responses up to 100x faster.

See Spice in action

Walk through your use case with an engineer and see how Spice handles federation, acceleration, and AI integration for production workloads.

Talk to an engineer