In our previous post, Databricks Zerobus — The Best Bus Is No Bus, we covered what Zerobus Ingest is, how it works under the hood, and why it’s a compelling alternative to Kafka when your sole destination is the Lakehouse. That post resonated — and the most common follow-up question I’ve gotten from principal architects and platform teams is:
“Great, but can I build a managed service on top of it? And how do I do transformations before the data hits the table?”
This post answers both questions in depth. If you’re a platform engineer, principal architect, or anyone designing a multi-tenant ingestion layer on Zerobus Ingest — this is for you.
The Two Questions Engineers Ask
When I sit down with customers who are evaluating Zerobus Ingest for production, the conversation almost always lands on two things:
“Can we do transformations on the fly?” — They want to scrub PII, normalize schemas, filter noise, or enrich records before they land in Delta.
“Can we build a managed service on top of Zerobus?” — They’re building a platform for internal teams or external tenants, and they want Zerobus as the ingestion backbone.
The short answer to both: yes, but not the way you might expect. Zerobus Ingest is intentionally simple — it does one thing extremely well (durable, high-throughput 12+ Gb/Second writes to Delta), and it does not try to be a transformation engine. That’s a feature, not a limitation.
Why Zerobus Doesn’t Do Transformations (And Why That’s Good)
Let me be direct: Zerobus Ingest does not offer native transformations.
This is a deliberate design choice. Zerobus is built in Rust for raw speed — WAL-durable ACKs in ~150ms, data queryable in Delta within ~5 seconds, 100 MB/s per stream, 10+ GB/s per table. Adding a transformation layer inside the ingestion path would compromise that performance and add complexity to a service whose entire value proposition is simplicity.
But here’s the thing: you don’t need Zerobus to do transformations. You need a tiered transformation strategy — and if you’re coming from a Kafka architecture, you probably already have the pieces in place.
The Tiered Transformation Architecture
This is the pattern I recommend to customers building on Zerobus.
The Four Tiers
TIER 1: PRODUCERS Mobile apps, web apps, IoT devices, APIs. Raw events: unvalidated, may contain PII, mixed formats.
TIER 2: YOUR INGESTION API SERVICE FastAPI / Flask / custom Go or Rust micro-service.
This is where in-flight transforms happen:
✅ PII scrubbing / anonymization
✅ IP → Geo mapping (database lookups)
✅ Compliance flags (privacy service calls)
✅ Format conversion (XML→JSON, CSV→JSON)
✅ Tenant identification & tagging
Latency budget: milliseconds per record. Then calls: stream.ingest_record_offset({cleaned_record})
TIER 3: ZEROBUS INGEST Databricks-managed, serverless.
Schema validation against Delta table
Durable ACK in ~150ms
Queryable in ~5 seconds
Auto-compaction via Predictive Optimization
TIER 4: LAKEHOUSE (Downstream)
Bronze → Silver: Deduplication, type casting
Silver → Gold: Multi-stage enrichment, joins
Complex aggregations, ML feature engineering
SCD Type 2, slowly changing dimensions
Tools: Lakeflow Declarative Pipelines / DLT / dbt
Why This Works
The key insight is that Tier 2 already exists in most architectures. If you’re currently using Kafka, you have an ingestion API service that receives events from producers and forwards them to Kafka. That same service can call the Zerobus SDK instead of the Kafka producer — and you can add your lightweight transforms there.
In-Flight (Tier 2): Record-level transforms via API calls to rules-based services and database lookups (IP-to-Geo mapping, privacy flags, anonymization)
Compliance & Privacy: Anonymization and pseudonymization happen before storage — critical for GDPR/CCPA
Downstream (Tier 4): Deduplication from Bronze to Silver, multi-stage enrichment from Silver to Gold
The beauty of this pattern is that your transformation logic is yours — you own it, you version it, you test it. Zerobus stays fast and simple.
Reference Implementation: Zerobus Ingest Station
Databricks has published a reference implementation called Zerobus Ingest Station — a FastAPI-based container service that wraps the Zerobus SDK to create customizable ingestion endpoints. This is exactly the Tier 2 pattern:
Define validation logic and transformations before pushing records to Zerobus
Ideal for external partners, IoT devices, or REST integrations
Deploy on Databricks Apps, Kubernetes, or any container runtime
Check it out: Zerobus Ingest Station on GitHub
Building a Multi-Tenant Managed Service
Now let’s talk about the second question: building a managed service. This is the pattern I’ve been recommending to healthcare, manufacturing, and SaaS platform customers.
The Architecture
Your managed service platform sits in front of Databricks:
Zerobus Supports Both Row-Level AND Bulk Ingestion
One misconception I keep hearing: “Zerobus is only for row-by-row ingestion.” Not true. Here’s the full picture:
Arrow Flight is the key for bulk scenarios:
Sends entire RecordBatch (thousands of rows) per call
Splits large batches into smaller transport messages automatically
Best when your producer already works with columnar data (Polars, DataFusion, pyarrow)
Supports ZSTD compression for improved throughput
If your managed service needs to handle both real-time events and periodic bulk loads, Zerobus has you covered with a single endpoint.
Regional Endpoints: There Is No Multi-Region Endpoint
This catches people off guard, so let me be clear: Zerobus endpoints are regional. There is no global or multi-region endpoint.
Your producers call the endpoint in the region where your workspace (and Delta tables) reside. Cross-region calls work but incur cloud provider egress charges. For best throughput, keep producer and endpoint in the same region.
What You Need to Call Zerobus
For your managed service’s Tier 2 layer, here are the parameters:
Critical permission gotcha: The service principal needs explicit grants — USE CATALOG, USE SCHEMA, SELECT, and MODIFY on the target table. ALL_PRIVILEGES does NOT work and will give you an invalid_authorization_details error.
The Honest Part: When NOT to Use This Pattern
Staying true to the spirit of my original post — here’s when this architecture doesn’t fit:
Customer Proof Points
This isn’t theoretical. Here are real customers stories:
Beusa Energy (Oil & Gas IoT): >99% cost reduction, ~3 seconds end-to-end latency, 22M+ rows/day from 6,000+ devices. “Zerobus Ingest saved us a lot of money, simplified our stack, and put us on a managed path forward.”
Joby Aviation (eVTOL Manufacturing): Streams gigabytes per minute from manufacturing sites. Reduced telemetry resolution latency from days to minutes. Custom on-premises forwarding agents push directly to Zerobus.
Toyota Motor Corporation (Factory IoT): Detects overheating factory conditions in minutes instead of hours. Uses Soracom Beam for global IoT connectivity → Zerobus → Delta.
Getting Started
If you want to build this pattern yourself:
Start with the docs: Zerobus Ingest Overview
Try the Zerobus Ingest Station: GitHub — a ready-made Tier 2 service
Read the deep dive: Deep Dive on Zerobus Ingest (Now GA)
See it at scale: Ingesting the Milky Way: Petabyte-Scale
Build an end-to-end app: Near Real-Time App with Zerobus + Lakebase
Conclusion
Building a managed service on Zerobus Ingest is not only possible — it’s the pattern Databricks designed for. Your service owns the API gateway, authentication, rate limiting, and lightweight transformations. Zerobus owns durable, high-throughput writes to governed Delta tables. The Lakehouse owns deep enrichment downstream.
The architecture is simple because Zerobus is simple. And sometimes the best architecture isn’t the most sophisticated — it’s the one that solves your problem.
Thanks for reading!
Jitesh Soni is a Specialist Solutions Architect at Databricks and the author of canadiandataguy.com. He writes about streaming, data engineering, and the things that break at 3 AM.






