Observability

Everything that touched your storage, and everything that stopped

The same access record that powers detection answers the operational questions. No instrumentation, no catalog access, no agents.

Read
GetObject
Write
PutObject
Delete
DeleteObject
Presigned
GET · presigned URL

Illustrative data.

What reCost keeps queryable

Data flow

Every request mapped from identity to bucket to prefix to operation, with volumes and direction across the estate.

Identity and object trace

Give it an identity and see everything it touched. Give it an object key and see everything that touched it.

Pipeline health

Write cadence learned per prefix, so a writer that goes silent surfaces on its own.

Asset inventory

Every bucket and prefix with owners, access classes, coverage state, dormant share, and storage class distribution.

Error tracking

Denials and failures by bucket, prefix, operation, identity, and client, classified by cause rather than counted.

Data lake access

Iceberg, Delta, and Hudi resolved from the object layer, including readers that skipped the catalog.

Coverage map

Which buckets are logged, which are not, and where other tooling is being denied.

What it looks like in your account

Illustrative data.

Monitored pipelines
3,350
1,040 healthy2,049 degraded242 stalled

Storage class distribution

180.7 TB · 915K objects
Standard
0.8 TB0%
Standard-IA
0.0 TB0%
Glacier
179.9 TB100%
Intelligent-Tiering
0.0 TB0%
Other
0.0 TB0%

Request profile

last 7 days
Requests (7d)717K
Succeeded / denied
100% ok
Operation mix
PUT 91%
Bytes served (7d)0 B
Active identities9
Access
100% ext

Identity relationships

last 30 days
IdentityClassRequestsFirst seenState
PE_Role
Machine
Application99% ok · 72M req≥90 daysestablished
svc:logdelivery…
Service
AWS Service100% ok · 5.8M req≥90 daysestablished
AmazonS3
Service
S3 Lifecycle100% ok · 2.7M req≥90 daysestablished
snowflake-role
Machine
Data Platform100% ok · 1.4M req≥90 daysestablished

Daily ingestion

last 14 days
normalspiketoday

Pipeline health

Job success does not mean data arrived. A Glue job reading an empty upstream partition completes cleanly and writes nothing. Firehose buffers records that never flush. reCost learns the write cadence of every prefix and pages when the silence exceeds it, with no instrumentation in the job.

AWS GlueFirehoseKinesisMSK ConnectSpark Structured StreamingFlinkAirflowdbt

Write cadence by prefix

last write per prefix, against learned cadence
PrefixWriterExpectedLast write
events-raw/dt=…firehose-clickstreamevery 5 min2 min agohealthy
warehouse/orders/glue-orders-nightlydaily 02:004 h agohealthy
exports/partner/msk-connect-s3hourly38 min agolate
warehouse/billing/glue-billing-etldaily 03:0017 d agostale
ml-features/v3/spark-feature-buildevery 6 h9 d agostale

2 prefixes stale beyond their inferred SLO

Illustrative data.

reCost data lake lineage, identity to client to asset

Data lake access

Iceberg, Delta, and Hudi tables all resolve to objects. reCost reconstructs table health from S3 Inventory and access log patterns, so it works across all three without engine-specific instrumentation, and shows which readers came through the catalog and which went straight to the files.

Snapshot counts, manifest bloat, orphaned files, small file accumulation, compaction lag, and checkpoint age, per table.

Four pipelines stopped writing. No alarm ever fired.

Job success does not mean data arrived. Write cadence per prefix is a direct measure of pipeline liveness, and it is already in your access logs.

Signals that exist nowhere else

Access

01Out-of-catalog reads. Which readers pulled Parquet or manifest files directly, bypassing the catalog entirely.

02External readers. Which identities outside your organization read from your buckets, what they read, and how much.

03Presigned URL consumers. The signing role appears in the log. The client that actually retrieved the object appears only in the user agent.

Operations

04Silent pipelines. Last write timestamp per prefix per writer identity, against a learned cadence.

05Broken integrations. Writes that authenticate but never land, so the pipeline reports success while the data does not exist.

06Coverage gaps. Which buckets have no access logging at all, and where your other tooling is being denied.

Tables and storage

07Iceberg. Snapshot count per table, manifest bloat, and whether expire_snapshots has ever run.

08Delta Lake. Log depth, checkpoint age, orphaned file volume, and small file ratios.

09Hudi. Compaction lag, timeline file counts, and log-to-base-file ratio per partition.

Observability, answered

No. Table health is derived from S3 Inventory and access log patterns, which is why it works across Iceberg, Delta and Hudi without engine-specific instrumentation.

Start with proof,
not a pitch.

Scoped read-only role, 30-day lookback, results in 48 hours.