In this article
Picture a data team that spent a quarter building a product-qualified lead model in Databricks. It scores every trial account each hour on teammate invites, feature adoption, and API calls. The lifecycle team hears about the top scorers on Monday, when someone exports a CSV, renames the columns, and uploads it to the messaging platform. By then a few of those accounts have already picked a plan, and a few more have gone quiet.
If your company runs on Databricks, your warehouse likely holds the most complete picture of each customer you have: product usage, billing, support history, and the outputs of every model your data team maintains. Your messaging platform often holds a thinner copy, refreshed on whatever schedule the last integration project allowed.
Customer.io now connects directly to Databricks as a reverse ETL source. You write SQL against a Databricks SQL warehouse, choose how often it runs, and each row becomes a profile update, an event, or an account record in Customer.io that can start or steer a journey. The PQL (product qualified lead) score your data team computes at 10:00 can be in a lifecycle marketer's segment a few minutes later, with no CSV in between.
More than 20,000 organizations use Databricks, including over 70% of the Fortune 500. And in Customer.io's 2026 customer messaging research, growth teams ranked data integration as a top priority, tied with scaling AI at 21%.
TLDR
- Customer.io's Databricks integration runs SQL queries against your Databricks SQL warehouse on a schedule, as often as every minute, and imports the results into Customer.io.
- Syncs can create or update people, events, custom objects (accounts, workspaces, teams), relationships, page views, screen views, and aliases.
- Customer.io connects through a Databricks service principal with read-only access, so your Unity Catalog grants decide exactly which tables it can read.
- Databricks
ARRAY,MAP, andSTRUCTcolumns arrive in Customer.io as structured JSON, ready to use in segments and message personalization. - More than 20,000 organizations rely on Databricks, which makes the warehouse the source of truth for customer data at a large share of companies.
What does the Customer.io Databricks integration do?
The Customer.io Databricks integration imports people, events, and objects from your Databricks SQL warehouse into Customer.io on a schedule you control. You connect a warehouse, write a query for each kind of data you want to bring over, and set a sync frequency. Customer.io runs the query, and each returned row becomes one operation in your workspace.
Each sync has a type that determines what those rows do.
- An
identifysync sets attributes on people, like plan tier, lifetime value, or a churn risk score. - A
tracksync records events, like "Trial Activated" or "Usage Limit Approaching," which can trigger journeys. - A
groupsync creates custom objects such as companies or workspaces, links people to them, and can store attributes on the relationship itself, like a person's role on an account. - Page and screen syncs bring in web and app views, and alias syncs merge identities that belong to the same person.
- A
{{last_sync_time}}placeholder lets each query read only rows that changed since the previous successful sync, which keeps runs small. - Semantic events let a
tracksync handle maintenance tasks, such as deleting a person, suppressing a person, or removing a relationship, when your warehouse records those changes.
Why your Databricks data should drive lifecycle messaging
The signals that predict conversion, expansion, and churn usually get computed in the warehouse, and a message can only respond to signals your messaging platform can see.
Think about where a definition like "active user" gets decided. Your data team writes it in Databricks, joining product events with billing and account data, and it feeds dashboards, forecasts, and board reports. If your messaging platform receives a separate event stream from the app, it ends up with its own version of "active," and your lifecycle reporting slowly drifts from the numbers leadership trusts. A direct sync means the segment in your marketing automation platform uses the same definition as the dashboard.
It also changes who waits on whom. Without a direct connection, every new attribute a marketer wants in a segment tends to become an engineering ticket: add a field to an export job, update a mapping, backfill history. With the Databricks integration, a data or marketing ops partner writes a SELECT statement, previews the results in Customer.io, and enables the sync. The lifecycle marketer builds the journey on top of it.
Governance stays where your data team already manages it. Customer.io authenticates as a service principal that needs only USE CATALOG, USE SCHEMA, and SELECT privileges on the tables you choose. If a table isn't granted, Customer.io can't read it.
How to use Databricks data for product-led growth
The Databricks integration supports product-led growth by letting you trigger lifecycle messages from the usage signals your data team already models in Databricks, such as activation milestones, product-qualified lead scores, and plan limits.
PLG signals are rarely single clicks. "Activated" might mean a workspace invited three teammates and connected its first integration within 14 days. "Product-qualified" might combine seat growth, feature breadth, and company size from an enrichment table. Client-side tracking can capture each piece, but the logic that turns those pieces into a decision lives in SQL. Syncing the result means your journeys act on the decision itself.
Here's how a few common PLG motions map to the integration:
PLG signal | Where it's computed in Databricks | What Customer.io does with it |
|---|---|---|
Activation milestone reached | Join of workspace, invite, and integration tables |
|
PQL score crosses a threshold | Scoring model output table |
|
Account nearing a plan limit | Daily usage aggregation against plan entitlements |
|
Admin vs. member roles | Account membership table |
|
For B2B and prosumer products, a group sync creates company or workspace objects in Customer.io with attributes like plan, seat count, and renewal date, then connects each person to their account. With relationship attributes, built in your query using Databricks' named_struct function, the same person can be an admin on one workspace and a viewer on another. Your journeys can then send billing and upgrade messages only to the people who can act on them, while everyone else on the account gets education about the features their team is using.
Imagine a collaboration tool's data team scores trial workspaces nightly. When a workspace's score passes 70, the identify sync updates the score on each admin's profile. A journey in Customer.io waits for that segment entry, sends an in-app message offering a guided upgrade, follows up by email two days later if nothing changes, and posts the account to the sales team's Slack channel through a webhook if the workspace has more than 20 seats. The model, the threshold, and the message sequence each live with the team best equipped to own them.
Can you trigger marketing campaigns from Databricks events?
Yes, through track syncs.
Each row your query returns becomes an event in Customer.io, and events can trigger journeys, update segments, and supply content for the messages themselves.
A basic event sync looks like this:
SELECT user_id AS userId,
event_name AS event,
products,
total_price AS value
FROM events
WHERE created_at >= timestamp_seconds({{last_sync_time}})
The WHERE clause matters. Comparing a timestamp column to {{last_sync_time}} means each run picks up only new rows, so a person who triggered an event once receives the journey once.
The more interesting use is computed events, the ones no single app action produces. Your data team can define "Engagement Dropped" as a week where a user's active days fall below half of their trailing four-week average, then emit that event from a scheduled query. A win-back journey starts from a signal the whole company agrees on, and the threshold can change in SQL without anyone rebuilding the campaign.
Event properties carry the context your messages need. Because ARRAY and STRUCT columns arrive as structured JSON, a recommendations model can write a list of suggested products or templates straight into the event, and your email can loop through that list with Liquid. The personalization comes from the model your data team already maintains, and the marketer controls how it shows up in the message.
Sync frequency is where timing and cost meet. You can sync as often as every minute, and each run wakes your SQL warehouse. The docs recommend a serverless warehouse with an auto-stop timeout so compute runs only while syncs do, and a frequency that matches how quickly the underlying signal changes: a nightly churn score doesn't need a one-minute sync, but an hourly usage-limit check might.
How do I connect Databricks to Customer.io?
You connect Databricks to Customer.io by creating a service principal in Databricks, granting it read access, and entering its credentials and warehouse details in Customer.io. The full walkthrough is in the Databricks integration docs. Here’s the high level:
- Create a service principal in your Databricks workspace and generate an OAuth secret. Save the application ID (client ID) and the secret.
- Grant read access with
USE CATALOGon the catalog,USE SCHEMAon the schema, andSELECTon each table you plan to sync, then give the service principal "Can use" permission on your SQL warehouse. - Allowlist Customer.io's IP addresses if your workspace uses IP access lists. The docs list addresses for US and EU regions.
- Connect in Customer.io by choosing the Databricks integration under Configure Data and entering the server hostname, HTTP path, client ID, client secret, catalog, and schema.
- Create your first sync by choosing a data type, writing the query, previewing results with Run Query, setting a frequency, and enabling it.
Most teams start with one identify sync for core profile attributes and one track sync for a single high-value event, then add syncs as journeys call for new data.
Setup using your AI agent in Customer.io
You can also set up syncs with an AI agent. If your team already uses one with access to your Databricks data model, point it at the Customer.io MCP: it brings the warehouse context, and the MCP supplies the Customer.io side, including the sync types, reserved columns, and the object types and attribute names already in your workspace. Your AI Agent in Customer.io can do the same from a plain-language request once your connection is in place.
Before anything runs on its own, it previews every sync against real sample rows, picks a schedule that fits around your other syncs, and keeps validation queries lean since each one runs on your SQL warehouse. After the first import, the Agent confirms the people, objects, and relationships landed in your workspace. Syncs that delete people get fully configured and stay switched off until you turn them on yourself.
What to look for in a marketing platform when you use Databricks
Databricks teams comparing marketing platforms should look at four things:
- What access the platform needs
- Whether you can sync with your own SQL
- Whether the data model fits how you describe customers
- How much work sits between a new signal and a live message.
For evaluation criteria beyond data, including channels, AI, and reporting, our customer engagement platform buyer's guide covers the full checklist.
A direct connection removes a layer of export jobs and mappings that someone has to maintain. Data model fit shows up in the details. If your warehouse models accounts, workspaces, and roles, your messaging platform needs somewhere to put them, and Customer.io's custom objects and relationships give accounts and roles a place to live alongside people. Customer.io also supports unlimited attributes, events, and custom objects, so a data team adding a new score doesn't have to negotiate for space.
Scale and support round out the picture once warehouse data starts triggering real journeys. Customer.io sent more than 100 billion messages in 2025, runs at 99.98% infrastructure uptime, and holds a 99% CSAT score with 24/5 global support.
A low-risk way to test the fit is to pick one signal your current setup can't act on today, like a PQL score or a computed churn event, sync it from Databricks, and build a single journey around it. That gives your data and lifecycle teams a concrete comparison using your own data.
Frequently asked questions
Does Customer.io integrate with Databricks?
Yes. Customer.io connects to Databricks as a reverse ETL source, importing people, events, objects, and relationships from a Databricks SQL warehouse using queries you write.
How often can Customer.io sync data from Databricks?
Syncs can run as often as every minute. Customer.io recommends a frequency that keeps runs from overlapping; if a sync is still running when the next one is scheduled, the next one is skipped and its data is caught up on the following run.
Does syncing from Databricks use Databricks compute?
Yes. Every sync and every query preview runs on a SQL warehouse in your Databricks account, and Databricks bills that compute. A serverless warehouse with an auto-stop timeout keeps costs predictable because it starts in seconds and shuts down between syncs.
Can Databricks data trigger journeys in Customer.io?
Yes. Events imported through a track sync can trigger journeys, and attributes imported through an identify sync can move people into segments that trigger journeys. For messages that need to fire within seconds of an in-app action, send those events directly through Customer.io's SDKs or API.
Which Databricks data types does Customer.io support?
Standard columns become attributes or event properties. ARRAY columns become array attributes, and STRUCT and MAP columns become object attributes, with no manual serialization required.
How does Customer.io keep Databricks data secure?
Customer.io connects through a Databricks service principal using OAuth client credentials. The service principal needs only read privileges on the catalogs, schemas, and tables you grant, and Customer.io stores the client secret securely and never displays it.
Put your Databricks data to work in Customer.io
If your customer data already lives in Databricks, the fastest way to see what the integration can do is to map one of your own signals to a journey. Book a demo to walk through it with our team, browse Customer.io integrations, or read what a customer engagement platform is for a broader look at how data and messaging fit together.






