Parquet: What floor are we standing on?
Row groups, column chunks, bloom filters, and the footer: a walk through the Parquet file format, with a browser tool to explore your own files.
Row groups, column chunks, bloom filters, and the footer: a walk through the Parquet file format, with a browser tool to explore your own files.
From raw Bluesky posts to Streamkap to derived tables and visualizations.
Register a Snowflake Horizon Iceberg catalog in oleander, then run SQL on the engine that fits the job — without moving data out of Horizon first.
Connect Neon, Supabase, or any Postgres. Query live through a SQL proxy with full governance, or sync tables into the lake once, hourly, or daily.
Install the MCP server for Claude and Codex from the official oleander plugin marketplace.
Hockey is shift-oriented, goalie-haunted chaos. A Celebrini shift, four attempts, one goal, and the expected-goals, RAPM, and WAR math that tries to make sense of it.
Managed Spark and lake are easy to start with and still grounded if you wish to bring your own infrastructure.
Exploring San Francisco 311 case data with oleander... From CSV ingestion to district demographics.
Managed Spark & Iceberg with OpenLineage by default.
The more I learn, the less I know.
Introducing Lake: write SQL with lineage captured automatically, collaborative public/private datasets, and thoughtful moderation.
Spark on EMR Serverless, write Iceberg tables in the Glue data catalog, parse JSON text, query with Athena, and monitor lineage & pipeline health in oleander.
Explore our journey building a browser-based Parquet viewer, from the initial implementation to our current DuckDB-powered solution with filtering and sorting capabilities.
A practical guide to implementing OpenLineage with Spark and Iceberg. Learn how to set up data lineage tracking in your data pipeline.
Oleander is the culmination of years of experience building data observability and metadata management tools. Learn about our journey from WeWork to OpenLineage.