AI Model Serving

Invoke and run AI models anywhere

Build faster, data-grounded AI without complex orchestration. Call, serve, and evaluate AI models locally or from hosted providers-directly in Spice.

feature_ai_model_serving_header

Do more with your data

Increase the reliability, performance, and security of data-driven AI workflows.

0x

up to 100x faster queries

0%

up to 80% cost savings on data lakehouse spend

0x

increase in data reliability for critical workloads

AI grounded in your data
Flexible models and deployments
Built-in tooling and memory
AI sandboxing
AI grounded in your data

AI grounded in your data

Run inference where your data lives. Spice removes external dependencies by colocating data, models, and compute in a single engine. Reduce latency and hallucinations by grounding models in federated and accelerated datasets.

View the docs

feature_ai_model_serving_data_grounding
Flexible models and deployments

Flexible models and deployments

Spice supports custom, open-source, and commercial models, adapting to your workload and regulatory requirements. Deploy and scale models anywhere-locally, in the cloud, or at the edge.

View the docs

feature_ai_model_serving_optionality
Built-in tooling and memory

Built-in tooling and memory

Augment models with native federation, search, and retrieval for more robust AI workflows. Enable long-term memory through persistent datasets so models can recall and learn across interactions.

View the docs

feature_ai_model_serving_built_in_tooling
AI sandboxing

AI sandboxing

Isolate, monitor, and control AI execution. Prevent unwanted model outputs and data leakage with sandbox environments for model inference and tool usage.

View the docs

feature_ai_model_serving_sandboxing

Integrations across all your data sources

Spice supports Hugging Face, OpenAI, Anthropic, xAI, NVIDIA NIM, and more out of the box. Serve and evaluate models across providers while keeping one consistent runtime for inference, evaluation, and observability.

platform_integrations

Deployed in production

Teams trust Spice to bring inference closer to their data, enabling low-latency, enterprise-grade AI across industries.

NRC Health logo
Basis Set Ventures logo
Tim Ottersburg

“Partnering with Spice AI has transformed how NRC Health delivers AI-driven insights. By unifying siloed data across systems, we accelerated AI feature development, reducing time-to-market from months to weeks - and sometimes days. With predictable costs and faster innovation, Spice isn't just solving some of our data and AI challenges - it's helping us redefine personalized healthcare.”

Tim Ottersburg

VP of Technology, NRC Health

Rachel Wong

“Spice AI grounds AI in our actual data, using SQL queries across many data sources. This brings accuracy to probabilistic AI systems, which are very prone to hallucinations.”

Rachel Wong

CTO, Basis Set

FAQs

Answers to common questions about serving AI models in Spice

Which AI model providers does Spice support?

Spice supports Hugging Face, OpenAI, Anthropic, xAI, and NVIDIA NIM out of the box, along with custom, open-source, and commercial models. Every provider runs behind one consistent runtime for inference, evaluation, and observability, so applications keep the same interface across providers.

Can I call AI models directly from SQL?

Yes. Spice supports LLM inference from SQL, so you can invoke large language models inside queries to summarize, translate, generate, or classify text alongside your data. Results return as query output, without additional orchestration code.

How does Spice reduce hallucinations in AI applications?

Spice grounds models in federated and accelerated datasets by colocating data, models, and compute in a single engine. Running inference where the data lives removes external dependencies and reduces both latency and hallucinations. The same foundation supports retrieval-augmented generation workloads.

Where can I deploy models served by Spice?

Models can be deployed and scaled locally, in the cloud, or at the edge. Spice adapts to your workload and regulatory requirements while keeping one runtime across each environment.

How does Spice secure model inference and tool usage?

Spice isolates, monitors, and controls AI execution with sandbox environments for model inference and tool usage. Sandboxing helps prevent unwanted model outputs and data leakage while keeping every interaction observable.

See Spice in action

Walk through your use case with an engineer and see how Spice handles federation, acceleration, and AI integration for production workloads.

Talk to an engineer