Data Playground

An interactive data-engineering lab for pipeline DAGs, data models, SQL analytics, graphs, and vector similarity.

Built a reproducible synthetic commerce pipeline with seeded event generation, deduplication, schema and lifecycle checks, and SQL analytics. Explore acquisition and retention scenarios, inspect quarantined records, and trace conversion, collected revenue, and cohort retention to the queries that compute them. The public React lab serves generated Python runs without a database; a bounded FastAPI service supports custom simulations when configured. A separate synthetic commerce dataset connects products, customers, and purchases in a relationship graph and explains cosine similarity over handcrafted product feature vectors. An engineering workbench exposes executable dependency DAGs, retry and failure traces, model grains and contracts, and architecture decisions grounded in the implementation. The original PostgreSQL and Streamlit experiment remains in the source repository.

All data is synthetic. Enable JavaScript to compare reproducible scenarios, inspect quality checks, and trace metrics to events and SQL.

Source and reproduction instructions · Back to portfolio