Catch broken pipelines and bad tables before queries break
S3-native data lake observability for Iceberg, Delta Lake, and Apache Hudi. reCost reads S3 access logs and S3 Inventory to score table health, catch stale writers, and attribute storage and query cost to the table that caused it. No agents, no catalog access.
Your queries are getting slower and you don't know why
Tables silently degrade
Snapshots accumulate, manifests bloat, small files multiply. By the time queries slow, you've been bleeding cost for weeks.
Pipelines fail invisibly
Glue jobs report success but write zero rows. Firehose stops. Nobody notices until an analyst opens a stale dashboard.
Cost is unattributable
You see $40K in Athena scans but can't tell which table, query, or team caused it.
"We had 42,015 snapshots on a single Iceberg table. Expiry had never run. Query planning was costing us on every scan."
Staff Data Engineer
Fintech, Series C
Three lenses. One S3 data source.
Iceberg, Delta, and Hudi health from metadata
- Snapshot growth, manifest-list size, small-files counts per table
- Detects when expire_snapshots, remove_orphan_files, rewrite_manifests, or VACUUM is overdue
- Orphan-file detection: finds Parquet files not referenced by any live snapshot, with wasted-storage dollar amount
Last-write tracking and silent-writer alerts
- Last-write timestamp per table, broken down by writer identity
- Supports Firehose, Kinesis, MSK, Glue, Spark Streaming, Flink, Airflow, dbt
- Pages you when a writer misses its SLO, before downstream consumers notice
Query observability across engines
- Athena, Trino, Glue, Spark, EMR, Databricks - scan size, GET requests, bytes read per query
- Per-table cost attribution mapped to writer, query, and team
- Filesize and partition-skew detection that flags compaction debt before it compounds
15.6 TB of Orphaned Files in S3, Invisible for 8 Months
How a Series C fintech discovered 15.6 TB of unreferenced Parquet files accumulating silently in their Apache Iceberg tables, and fixed it in a day.
Stop flying blind on your data lake
5-minute setup. No agents. Works with your existing AWS stack.
Book a Demo