Chat with your Databricks data in plain English
Updated 2026-08-25
An AI data analyst is most useful when it runs on your real data. This guide shows how to connect Databricks so you can ask questions in plain English and get charts, dashboards and reports back — without writing SQL by hand.
The setup takes a few minutes: connect with a read-only user, let the analyst read the schema, and ask your first question. Your credentials are encrypted, and the analysis runs on your own AI provider key.
- 1
Get your Databricks connection details
You need the SQL warehouse's server hostname and HTTP path, plus a personal access token. Find these under the SQL warehouse's connection details in Databricks.
- 2
Create a read-only user (recommended)
Grant the token's user SELECT on the catalogs and schemas you want analysed via Unity Catalog. Keep it to read privileges only.
- 3
Create your account and add your AI key
Sign up, then add your own AI provider key (bring-your-own-key). If you don't have one yet, a free key takes about a minute to create — your key is encrypted and used only for your requests.
- 4
Add the Databricks connection
Add a data source, choose Databricks, and enter the server hostname, HTTP path and access token. Click Connect: it saves the token encrypted and then immediately calls Databricks, so you get either “Connected — N tables found” or the exact error Databricks sent back. A source that fails is still saved, but Databricks details cannot be edited in place — remove it and add it again with corrected values.
- 5
Let the AI read and annotate your schema
The analyst annotates your Unity Catalog tables and columns. With dbt on Databricks, point it at your target schema so it uses your models.
- 6
Ask your first question in plain English
Ask something like "Weekly active users by product surface". On a question the analyst judges complex — a join, a query spanning two sources, several steps — it shows a readable plan and waits for you to approve it before running anything; on one it judges simple it runs first and shows you afterwards. Either way you get the same two controls on every answer: Show SQL gives you the exact query that produced the number, and Show table gives you the rows it returned with their count. That is how you check a figure instead of trusting it. Refine by chatting, then pin it to a dashboard or save it as a report.
Why connect Databricks instead of exporting
A static export goes stale the moment you download it, and when an assistant only ever sees that exported file it has to guess what your columns mean. A connected AI data analyst reads live Databricks data, shows you a readable plan before running anything on the complex questions (simple ones it runs first and shows you after), and saves what your fields mean against the source itself — so the next question already knows them, and answers stay current and consistent.
Databricks: things to know
- Use a running SQL warehouse; a stopped warehouse takes a moment to resume on the first query.
- The HTTP path is specific to each SQL warehouse — copy it from that warehouse's connection details.
Example questions to ask your Databricks data
- Weekly active users by product surface
- Revenue by segment this quarter
- Model inference volume by day
Keeping it safe
- Connect with a read-only user. That grant is what actually stops a write, and it is yours to set — and when you connect we check what that account is allowed to do, and show you the answer.
- On our side every query is checked before it runs and refused if it contains a write statement — DROP, DELETE, UPDATE, INSERT and fourteen others, eighteen keywords in all. It is a keyword guard, not a full SQL parser, so it can only refuse what it can spell — which is why the read-only grant on your side is the control that actually holds. The two limit different things and neither replaces the other: your grant decides what we can be given, our guard decides what we can send.
- Connection details and your AI key are encrypted at rest.
- We train no model on your data — we do not have one; the analysis runs on your own provider key, so the terms that decide what happens to what you send are your provider's, not ours. We read the tables an answer needs into memory for that session rather than keeping a mirror of your database, and the rows an answer returns stay in your own account until you delete them.
Frequently asked questions
Use a SQL warehouse — copy its server hostname and HTTP path from the connection details.
Through Unity Catalog grants on the token's user; grant read-only SELECT on the schemas you want analysed.
Related
Get your end-to-end AI business intelligence now.
Ask in plain English, pin the answer to a dashboard, export it as CSV, Excel, PDF or PowerPoint — on the free tier that whole loop runs on one database source plus Google Sheets, and the field definitions you save carry into every later session, including after a schema refresh. One exception you should hear from us rather than discover: on a Google Sheet, re-running the tab picker — even only to add a tab — clears the definitions saved for that source and reads its schema fresh. Connecting a database is the technical part: a read-only user and network access, once. If your numbers live in a spreadsheet, none of that applies: put the file in a Google Sheet and it connects on the free tier, with no read-only user and no open port. Uploading a CSV or Excel file directly is a Pro feature, and every new account gets 14 days of Pro with no card. The one thing no plan waives is your own AI provider key, and a Google Gemini key is free to create.
Every new account starts on a 14-day Pro trial with no card. After it lapses, the free tier stays free.