Monorepo for Databricks Zerobus Ingest SDKs.
GA: This SDK is generally available and supported for production use cases. Minor and patch version updates will not contain breaking changes. Major version updates may include breaking changes.
We are keen to hear feedback from you. Please file issues, and we will address them.
Zerobus is a high-throughput streaming service for direct data ingestion into Databricks Delta tables, optimized for real-time data pipelines and high-volume workloads.
| Language | Directory | Package |
|---|---|---|
| Rust | rust/ |
databricks-zerobus-ingest-sdk |
| Python | python/ |
databricks-zerobus-ingest-sdk |
| Go | go/ |
github.com/databricks/zerobus-sdk/go |
| TypeScript | typescript/ |
@databricks/zerobus-ingest-sdk |
| Java | java/ |
com.databricks:zerobus-ingest-sdk |
| C++ | cpp/ |
Source / CMake (zerobus::zerobus) |
| C# | dotnet/ |
Databricks.Zerobus.Ingest.Sdk |
We try to provide prebuilt native binaries for the following platforms:
| Platform | Architecture |
|---|---|
| Linux | x86_64 |
| Linux | aarch64 |
| Windows | x86_64 |
| macOS | x86_64 |
| macOS | aarch64 (Apple Silicon) |
Note: We do not currently have macOS CI runners, so macOS binaries are built locally and may not be available for every SDK or release. If your platform is not supported or you encounter compatibility issues, you can build from source or file an issue.
Before using any SDK, you need the following:
After logging into your Databricks workspace, look at the browser URL:
https://<databricks-instance>.cloud.databricks.com/o=<workspace-id>
- Workspace URL: The part before
/o=(e.g.,https://dbc-a1b2c3d4-e5f6.cloud.databricks.com) - Workspace ID: The part after
/o=(e.g.,1234567890123456)
Note: The examples above show AWS endpoints (
.cloud.databricks.com). For Azure deployments, the workspace URL will behttps://<databricks-instance>.azuredatabricks.net.
Create a table using Databricks SQL:
CREATE TABLE <catalog_name>.default.<table_name> (
device_name STRING,
temp INT,
humidity BIGINT
)
USING DELTA;Replace <catalog_name> with your catalog name (e.g., main).
- Navigate to Settings > Identity and Access in your Databricks workspace
- Click Service principals and create a new service principal
- Generate a new secret for the service principal and save it securely
- Grant the following permissions:
USE_CATALOGon the catalog (e.g.,main)USE_SCHEMAon the schema (e.g.,default)MODIFYandSELECTon the table
Grant permissions using SQL:
-- Grant catalog permission
GRANT USE CATALOG ON CATALOG <catalog_name> TO `<service-principal-application-id>`;
-- Grant schema permission
GRANT USE SCHEMA ON SCHEMA <catalog_name>.default TO `<service-principal-application-id>`;
-- Grant table permissions
GRANT SELECT, MODIFY ON TABLE <catalog_name>.default.<table_name> TO `<service-principal-application-id>`;The service principal's Application ID is your OAuth Client ID, and the generated secret is your Client Secret.
Pick the record format that matches your data.
- JSON: schema-free ingestion. Pass a JSON string or a native object (dict, map, and so on) and the SDK serializes it. No compilation step. Good for getting started or dynamic schemas.
- Protocol Buffers: strongly-typed, schema-validated ingestion. More compact on the wire than JSON. A typical choice for production workloads that are not already producing Arrow.
- Arrow Flight: Apache Arrow
RecordBatchdata over the Arrow Flight protocol. Best when the workload is columnar or batched, or the application already produces Arrow (pyarrow, arrow-rs, DataFusion, Polars).
JSON and Protocol Buffers share one stream API, available in every SDK. Arrow Flight is a separate columnar API, available in the SDKs listed below.
| SDK | JSON / Protobuf | Arrow Flight |
|---|---|---|
| Rust | Available | Available since 2.8.0 |
| Python | Available | Available since 1.8.0 |
| Go (cgo) | Available | Available since 1.6.0 |
| Pure Go | Available | Not available |
| TypeScript | Available | Available since 1.3.0 |
| Java | Available | Available since 1.6.0 |
| C++ | Available | Available since 0.3.0 |
| .NET (C#) | Available | Not available |
Records are sent as JSON or Protocol Buffers on the same stream API.
For Protocol Buffers, use proto2 syntax with optional fields so nullable Delta table columns are represented correctly. Instead of writing .proto files by hand, each SDK ships a tool that generates a protobuf schema from an existing Unity Catalog table. See the individual SDK READMEs for language-specific usage.
Send Apache Arrow RecordBatch data directly to Zerobus. A good fit when:
- The workload is naturally columnar or batched — analytics pipelines, gateways aggregating short windows of rows, wide or numeric schemas where row-by-row serialization adds noticeable CPU overhead.
- The application already produces Arrow data — pyarrow, the arrow-rs crates, DataFusion, Polars, or other libraries built on Arrow.
For sparse, one-row-at-a-time traffic, JSON or Protocol Buffers are usually simpler. Most SDKs that expose Arrow Flight ship a runnable examples/arrow/ directory; see each SDK's README for details.
Delta column types map to Arrow and proto2 as shown below. JSON is not a typed mapping: objects use the same logical types as the table, and the SDK serializes them without a compiled schema.
For Protocol Buffers, declare fields optional when the Delta column is nullable.
| Delta Type | Arrow Type | Proto2 Type |
|---|---|---|
| TINYINT, BYTE | Int8 |
int32 |
| SMALLINT, SHORT | Int16 |
int32 |
| INT | Int32 |
int32 |
| BIGINT, LONG | Int64 |
int64 |
| FLOAT | Float32 |
float |
| DOUBLE | Float64 |
double |
| STRING, VARCHAR | LargeUtf8 |
string |
| BOOLEAN | Boolean |
bool |
| BINARY | LargeBinary |
bytes |
| DECIMAL | LargeUtf8 |
string |
| DATE | Date32 |
int32 (days since Unix epoch) |
| TIMESTAMP | Timestamp(Microsecond, UTC) |
int64 (microseconds since Unix epoch) |
| TIMESTAMP_NTZ | Timestamp(Microsecond) (no timezone) |
int64 (microseconds since Unix epoch) |
| ARRAY<type> | List (item field item) |
repeated type |
| MAP<key, value> | Map (entries field entries with keys and values) |
map<key, value> |
| STRUCT<fields> | nested struct | nested message |
| VARIANT | struct of metadata and value, both non-null LargeBinary |
string (JSON) |
DECIMAL is encoded as text on both paths today (LargeUtf8 / string). STRING maps to Arrow LargeUtf8, not Utf8.
Ingestion is asynchronous in every SDK. An ingest call returns as soon as the record is queued — the SDK sends it and tracks its acknowledgment on a background task. To confirm that records were durably committed, call flush(); it returns once everything queued so far has been acknowledged.
The idiomatic flow is therefore ingest in a loop, then flush() — once at the end of a bounded batch, or periodically for a long-running stream. Where the SDK supports it, you can instead register an ack callback and be notified as records commit, without blocking at all.
Each ingest also returns the record's offset, and wait_for_offset(offset) blocks until that offset is acknowledged. That's useful when a particular record must be confirmed before you continue; because acknowledgments are ordered, waiting on the last offset of a run confirms the whole run. The one thing to avoid is waiting on every record inside a tight loop — that turns the asynchronous pipeline into a synchronous request/response and limits throughput to a single record per network round-trip.
See each SDK's README for exact method names and a runnable example.
All SDKs support HTTP CONNECT proxies via environment variables, following gRPC core conventions. The first variable found (in order) is used:
| Proxy | No-proxy |
|---|---|
grpc_proxy / GRPC_PROXY |
no_grpc_proxy / NO_GRPC_PROXY |
https_proxy / HTTPS_PROXY |
no_proxy / NO_PROXY |
http_proxy / HTTP_PROXY |
The no_proxy value is a comma-separated list of hostnames (suffix-matched) or * to bypass the proxy entirely.
export https_proxy=http://my-proxy:8080
export no_proxy=localhost,127.0.0.1The SDK establishes a plaintext HTTP CONNECT tunnel through the proxy, then performs a TLS handshake end-to-end with the Databricks server. The proxy never sees decrypted traffic.
See CONTRIBUTING.md. Each SDK also has its own contributing guide with language-specific setup instructions.
This project is licensed under the Apache License 2.0. See LICENSE for the full text.