Arrow-native data movement.
Engineered for scale.

Fast, resilient, and transparent replication between production databases and data warehouses. Zero serialization tax. Zero dirty state.

EttaflowEttaflow ETL/dashboard

System Overview

Monitor real-time replication stats, pipeline health, and connector activity.

PIPELINES
5

4 active syncs →

CONNECTIONS
5

Sources & destinations →

DATA VOLUME SYNCED
4.82 GB

3,842,910 rows →

SYNC SUCCESS RATE
100%

49 of 49 runs succeeded →

Recent Executions

Latest system-wide synchronization runs. Click row for detail.

PipelineStatusData SyncedDurationStarted
pg-to-clickhouse Completed412.8 MB8.4s06:14:12
mysql-to-snowflake Completed184.2 MB4.1s06:12:00
events-to-bigquery Streaming92.5 MBIn progress06:15:02
billing-batch-sync Completed64.1 MB2.8s06:10:45
pg-to-clickhouse Completed512.0 MB10.2s06:05:30

Data Replication Share

Cumulative volume per pipeline.

pg-to-clickhouse2.41 GB (50%)
mysql-to-snowflake1.18 GB (25%)
events-to-bigquery850.4 MB (18%)
billing-batch-sync320.1 MB (7%)
Arrow Columnar In-Memory Speed
No Row-Tax (MAR) Pricing
Partition-Parallel Concurrency
Zero Dirty State Rollback
WAL-Safe Object Storage CDC Buffering
Granular VPC Compute Sizing
Stateful Auto-Resume
Granular Schema Evolution

Engineered for throughput. Built for reliability.

Arrow-native streaming pipelines with hybrid VPC execution and zero dirty state guarantees.

public.events_stream
"id" uuid PRIMARY KEY,
"created_at" timestamp DEFAULT now(),
+ "vector_embedding" vector(1536) Added v1.5
~ "status" varchar(64) → enum Evolved

Schema Evolution & Metadata

Automated schema discovery, auditable DDL versioning, and non-blocking column additions across active CDC pipelines.

Sync Types
Full Refresh Incremental CDC
Sync Modes
Append Upsert

Sync Types & Modes

Full Refresh, Incremental, and CDC sync types with Append or Upsert destination strategies.

Object Storage Buffer

WAL-safe fallback staging to cloud object storage during destination throttling.

Control Plane Ettaflow Cloud Orchestrator API
User Private VPC Data Plane Workers Autoscaled CPU/RAM
HYBRID DEPLOYMENT

Hybrid VPC Compute & Concurrency

Control plane runs in Ettaflow Cloud while autoscaled data plane workers execute inside your private VPC. Configure exact CPU and memory caps per pipeline for cost and throughput control.

[P1] CTID
[P2] CTID
[P3] CTID

Partition-Parallel Speed

Parallel initial loads and CTID scans across large partitioned tables.

Stateful Auto-Recovery

Stateful auto-resume from exact failure boundary after network interruptions.

Commit Ledger 100% Exactly-Once

Zero Dirty State Rollback

Transactional commit ledgers undo uncommitted changes automatically on failure.

Supported Connectors

Native columnar drivers built for zero-copy Arrow memory execution.

PostgreSQL

MySQL

Snowflake

BigQuery

ClickHouse

More Connectors

* Our connector ecosystem is continuously expanding. Additional databases, warehouses, Message queues, SaaS and API connectors and open data lakehouse formats are added on an ongoing basis.

Founder Manifesto
“High-scale data movement shouldn’t punish your engineering budget.”

We built Ettaflow after years of running large-scale data infrastructure. Engineering teams were trapped between two extremes: expensive SaaS vendors charging punitive row taxes as data grows, or open-source tools burning massive compute on slow row-by-row streaming. Ettaflow is built on modern distributed systems principles: Arrow-native memory efficiency, partitioned concurrency, and transparent cost governance.

100% Open Source (Apache 2.0) Built in Go + Arrow + Temporal $0 MAR Row Tax
Vishal Goswami

Vishal Goswami

Co-Founder

LinkedIn
Shubham Pandey

Shubham Pandey

Co-Founder

LinkedIn

Transparent open source evolution.

Engineered completely in the open. Track our development milestones from core replication to AI agent context and enterprise governance.

ARROW MEMORY STREAM NODEActive Focus
PHASE 01

Core Execution Engine

High-throughput Arrow memory execution engine with production database drivers.

  • Zero-copy Arrow memory engine & ADBC transfer
  • 5 Core launch connectors: PostgreSQL, MySQL, Snowflake, BigQuery, ClickHouse
  • VPC private worker deployment & auto-concurrency
  • Stateful crash recovery & zero dirty state rollback
MCP SERVER & AI CONTEXT PROTOCOLIn Planning
PHASE 02

Ecosystem & AI Agent Context

Expanded format drivers and semantic schema context for AI coding agents.

  • Connectors for Apache Iceberg, Parquet, Kafka, MongoDB, MS SQL and more
  • Model Context Protocol (MCP) server & CLI for AI coding agents
  • Schema intelligence, versioned change logs & column annotations
  • Declarative SQL replication & native dbt execution
DATA QUALITY GATES & RBAC SECURITYFuture Vision
PHASE 03

Data Quality & Governance

In-flight data quality gates and enterprise compliance audit trails.

  • Automated quality assertions & schema drift protection
  • Enterprise RBAC, audit trails & webhook alerting
  • Multi-region failover & policy-based encryption

Get Early Access & Book a Demo

Submit your workload details below, schedule an instant 1-on-1 session via Cal.com, or reach out directly to our founding engineering team.

Apply for Beta & Demo

Tell us about your pipeline requirements to get priority access and custom benchmark sizing.

No sales spam. Direct access to founding engineers.

Schedule an Instant Call

30-Min Intro Session

Book on Cal.com
✓ Direct Cal.com sync✓ No sales pitch

Write to us instead

vishal@ettaflow.io
Send
shubham@ettaflow.io
Send