Eclusa, by Aule

v2.0

Install it in your workspace, run setup and register one configuration row per dataset. Eclusa handles loading, quality and auditing without requiring a new pipeline for every source.

Requires a Databricks workspace with Unity Catalog enabled. Nothing runs outside it.

eclusa.whl

installed on the cluster, no external service

installs
control catalogueone row per dataset
source.format
csv
load.strategy
merge
quality.severity
fail
validates the contract
qualitya blocking rule stops the write
writes
bronzefaithful source, audited per run01
silvertyped contract and quality applied02
goldconsumption layer, modelled03

The repeated foundation

The first months of a data platform are often consumed by necessary but repetitive work.

One pipeline per source

CSV, ERP, API and spreadsheets become separate implementations, with standards that drift across teams.

Scattered quality rules

Rules live across notebooks and may detect a problem only after the data has already been written.

Runs without shared context

Logs, history and load identifiers vary by project, making failures harder to trace.

How it works

The flow is parameter-driven: configuration is validated before execution and behavior stays consistent across every dataset.

  1. 01 / 04

    Install

    Add the Python wheel to the customer's Databricks workspace.

  2. 02 / 04

    Run setup

    The command provisions the control catalog in Unity Catalog. It is idempotent and safe to run again.

  3. 03 / 04

    Register the dataset

    Define source, destination, load strategy and quality rules in a typed contract.

  4. 04 / 04

    Execute

    The engine validates, ingests and records the run with quality and auditing inside the customer's environment.

Available today

design partners

Unity Catalog setup

Managed tables with no mounts, fixed paths or external infrastructure.

Typed contract

Versioned configuration with fail-fast validation before execution.

Quality before write

Warn or fail rules. A blocking violation keeps bad data from being written.

Run-level auditing

Every load gets a correlated run_id recorded in the customer's environment.

Adaptable conventions

The default profile is medallion, while naming can follow the customer's standards.

Operational CLI

Setup, validation, registration and execution through clear, reproducible commands.

9 source formats

Current batch engine support, extensible through plugins without changing the core.

CSVJSONParquetDeltaAvroXMLORCTextExcel

4 load strategies

The strategy is declared per dataset and applied consistently.

  • Overwrite
  • Append with deduplication
  • Idempotent key-based merge
  • Atomic replace-where

Control by architecture

The product is installed in the customer's Databricks workspace. Aule does not host, access or store the processed data.

Data in your environment

Configuration, execution and auditing remain inside the customer's cloud and policies.

No external service

Setup uses Unity Catalog tables and does not depend on a vendor database or infrastructure.

Technical questions