Unstable production workloads
Impact: Failed jobs, delayed pipelines and unclear recovery ownership affect reporting and downstream operations.
Response: Review orchestration, retries, dependencies, monitoring, runbooks, escalation and resilience practices.
Performance without a clear diagnosis
Impact: Slow queries and inefficient workloads lead teams to add capacity without understanding root causes.
Response: Examine query patterns, file layout, compute choices, concurrency, caching and workload design.
Growing platform spend
Impact: DBU and cloud costs rise without reliable allocation, accountability or optimisation priorities.
Response: Review utilisation, policies, idle resources, workload schedules, tagging and cost attribution.
Fragmented governance
Impact: Inconsistent ownership, permissions and catalogue structures reduce trust and increase control effort.
Response: Assess Unity Catalog design, access models, ownership, lineage and governance workflows.
Security and audit concerns
Impact: Privileged access, secrets, logging, sharing and network gaps create avoidable exposure.
Response: Review control design, evidence, exceptions, responsibilities and specialist review points.
Operational knowledge concentrated in a few people
Impact: Support quality and decision speed depend on undocumented individual knowledge.
Response: Examine runbooks, backup staffing, change control, incident processes and knowledge transfer.