
// DataSpace
DataSpace is an all-in-one, AI-assisted data platform: ingest data, create pipelines, write transformations, preview datasets, monitor workloads, and seamlessly generate charts and dashboards.
Write arbitrary Python code in an integrated IDE. Transformations run in secure Docker containers, leveraging the speed of Polars and the efficiency of Parquet files.
Develop in isolated branches and merge logic seamlessly.
Transformations run securely in isolated environments.
Lightning fast analytical queries backed by Parquet.

Generate and publish dashboards and charts. Handle unstructured data like documents and images with artifact storage. Time-travel through dataset versions by time or branch.
Visualise data & plots
Unstructured Data
DataSpace analyses source code to generate column-level lineage in real-time. Build visibility allows you to track the evolution of row counts, file sizes, and duration over time.
Resolve upstream and downstream dependencies instantly.
Monitor dataset evolution (rows, size, duration) per build.
Click to explore the pipeline lineage and execution states.
Pipeline structure defined on main branch. All nodes initialized in neutral state.
Scheduled execution triggered. Ingest and Join processing completed successfully.
A failure occured during a build
Introduced a "Sanitize" node to filter bad data before validation to fix the build failure.
Running sanitation checks on the new node.
Feature branch build passed all checks. Sanitize node successfully filtered outliers.
Merged feat/sanitization into main. Production pipeline updated with robustness fixes.
Explore, analyze, and visualize your data using natural language directly on top of your existing workspaces. No data leaves your infrastructure.
Iterative data exploration. Automatically reads workspace README.mds to learn your business logic.
Generates charts and data tables instantly.
Fully offline operation. 100% data sovereignty.
DataSpace allows for complete self-hosting. Segregate projects into workspaces and control access with resource-based permissions.
Deploy on-premise or in your private cloud. Keep total control over your data.
Resource-based access control for granular permissions.
Strict segregation of projects, data, and pipelines into workspaces.
////////////////////////
//////////////////// 09
Strict segregation of projects. Manage resources, access control, and configurations independently for each workspace.
Blazing fast transformation engine using Polars and efficient Parquet file storage for high-performance ETL.
Full versioning support. Branching, commits, and time-travel for datasets ensuring total reproducibility.
Code-based lineage analysis. Visualize dependencies at the column level to understand data flow and impact.
Manage unstructured data. Store and analyze documents, images, and arbitrary binaries within your data pipeline.
Automate pipelines. Schedule builds based on time or events, with full visibility into duration and statistics.
Deploy DataSpace natively within your own infrastructure. Total control over your network and hardware.
Declarative health checks and data quality assertions that run concurrently with your pipelines.
Chat with your data using natural language. Fast, secure, and operates entirely on your infrastructure.