Executive Summary

Cloudera published data readiness results for the energy and utilities sector on September 17. Of the respondents, 86% report visibility into where their data resides, and 79% say they can access that data regardless of its format or location. This article looks at what those shares actually measure, and at what changes once they stand next to the other documents Cloudera has published from the same survey.

The number that catches the eye is 65%. The release calls it the share reporting that all or nearly all of their data is governed. Yet a piece Cloudera wrote itself in April about this same survey says that once the bar narrows to all, only two sectors clear 20%, telecommunications and software and technology. Energy and utilities is neither of them. What separates 65% from under 20% is a single word, nearly.

Sections 1 through 3 follow what Cloudera's own documents say about this survey: the April and September releases, the sector by sector announcements, the company blog, a piece published under its own name on Forbes, and the landing page for the telecom report. Section 4, where the question turns to what a team can put in place of a yes, is this article's reading and is not in any of those documents.

Key Figures

Source: Cloudera energy and utilities release (2026-09-17) · Global release (2026-04-14)

65%

Report all or nearly all data governed

Energy and utilities respondents. On the all-only bar in the same survey, no sector besides telecom at 33% and software at 26% clears 20%

84% ↔ 18%

Confidence set beside full governance

Across the full sample, 84% expressed confidence in their own data, while the share reporting full governance was 18%

54% ↔ 89%

Two values one index gives for telecom

Data visibility runs 54% in the April release and 89% on the telecom report page. The four sector cuts cluster between 82% and 89%

25%

Cost overruns lead the reasons returns fall short

Cost overruns lead in energy and utilities alone. Across the full sample data quality led at 22%, with cost overruns at 16%

1

What the September 17 Release Says, and What It Leaves Out

The release is headlined on energy and utilities leaders looking beyond AI adoption to build trusted data foundations. Three shares carry the argument in the body. 86% of respondents report visibility into where their data resides, 79% say they can access their organization's data regardless of format or location, and 65% report that all or nearly all of their data is governed. Cloudera reads that last item as reflecting continued investment in data management and modernization initiatives.

The tone turns in the paragraph that follows. For 25% of energy and utilities respondents, cost overruns emerged as the leading reason AI and analytics investments fail to deliver expected returns. Workforce problems come up alongside that. Data literacy and training are named as obstacles to effective data use, and no share is attached to them. One sentence attempts a comparison between sectors: compared to other sectors surveyed, energy and utilities organizations reported relatively few operational disruptions caused by infrastructure performance. That sentence carries no number either.

So this one release on its own does not give a reader the methodology. How many people answered, which countries they sit in, how the questions were worded, how many points the answer scale had: none of it is on the page. In place of a link to the report body there is an invitation to visit Cloudera's website.

It would be wrong, even so, to say the methodology is nowhere. This announcement is a sector cut of the Data Readiness Index 2026, first published on April 14, and the original release does state the method. Researchscape fielded the survey among 1,270 IT leaders at companies with more than 1,000 employees, from January 22 to March 3, across the AMER, EMEA and APAC regions, and the results were weighted to be representative of the overall GDP of surveyed countries.

What is missing is the breakdown of those 1,270 by sector. Neither the September release nor the April one says how many energy and utilities respondents this survey had. Country distribution is not published at the sector level either. The confidence interval for the survey as a whole can be estimated, then, while the number of answers behind the 86%, 79% and 65% quoted this time stays out of a reader's reach. A 65% drawn from 100 people and a 65% drawn from 30 are different numbers, and the documents hold nothing that would tell them apart.

This is not the first sector cut either. Cloudera pulled telecommunications out of the same survey and opened a report landing page for it, published healthcare results on June 24 and manufacturing results on September 8. Energy and utilities followed on September 17. All four are cuts of the same 1,270 respondents with one sector lifted out, and not one of the four documents records how many respondents that sector had.

Four sector cuts pulled from the same 1,270 respondents Full sample 1,270 IT leaders Apr 14 Full results (method disclosed) Jun 24 Healthcare Sep 8 Manufacturing Sep 17 Energy & utilities n undisclosed n undisclosed n undisclosed All four cuts come from the same 1,270 respondents with one sector pulled out. None of the four documents states that sector's respondent count — a 100-person 65% and a 30-person one look alike.
▲ Original Pebblous diagram — four sector cuts from the same survey and their undisclosed sample sizes

One thing is worth stating plainly. This survey is not an audit. Nobody opened each company's data catalog and access logs to check. These are the shares of IT leaders who answered yes about their own organizations. The 86% is not the fact that data locations are known; it is the proportion of people who think they know.

2

Which Question Does the 65% Answer?

The release records the question behind the 65% as all or nearly all of their data is governed. Two bars are bundled into one phrase. How much that bundle covers is visible in a piece Cloudera itself put on Forbes BrandVoice on April 14. That piece is sponsored content written by Cloudera under its own name, so the figures below are not a critic's recalculation but the account of the party that commissioned the survey.

On the wide bar of all or nearly all, the public sector stands highest at 72%, with healthcare at 66%, energy and utilities at 65% and software and technology at 63%. When only the respondents who answered all are counted, both the ranking and the size change. Telecommunications comes highest at 33%, software and technology follows at 26%, and no other sector reaches 20%. Energy and utilities is neither telecommunications nor software and technology, so the share in this sector reporting that all of its data is governed sits below 20%.

This does not rest on the sponsored piece alone. Cloudera's landing page for the telecom report also states that only 33% of telecom organizations report that all of their data is fully governed. That is the same value the BrandVoice piece gives as the highest under the all bar. The claim that the top sector on the strict bar stops at 33% therefore spans two separate Cloudera documents.

The same thing happens across the full sample. In the April release, fewer than one in five respondents, 18%, said their data was fully governed. With the question set at most of their data, the share becomes 71%. And 84% of respondents felt confident in the accuracy, completeness and alignment of their organization's data. The same people, inside the same questionnaire, reported 84% confidence and 18% full governance.

How the figures move when the bar changes inside one survey 0% 50% 100% All respondents Confident in their data 84% Energy and utilities All or nearly all governed 65% Energy and utilities All data governed Not disclosed · under 20% Telecom All data governed 33% Software and tech All data governed 26% All respondents All data governed 18% The orange bar is the figure quoted in the September release. The dashed box is where this sector lands once the bar narrows to all data. Only telecom and software/tech cleared 20%, which puts energy and utilities below both.
▲ Pebblous original diagram — figures from: Cloudera energy and utilities release, April global release, Cloudera BrandVoice

This gap can be read as organizational immaturity. Several reports on the survey did settle on that reading, with confidence running ahead of control. It holds up. There is a second reading as well. In answers where an organization grades itself, a question that turns slightly more generous lifts the share a long way. The words nearly all let a respondent forgive the unfinished parts of their own organization, and the word all takes that room away. The distance between 65% and under 20% is the size of that room.

None of this makes 65% a wrong number. It is faithful to the question. The trouble is that the question it answers and the question a reader has in mind while looking at it are not the same one. The reader takes it to mean that roughly six and a half companies in ten have their data under management, while the survey asked whether respondents judged their own organization to fall under all or nearly all.

3

The Same Item, Measured With a Different Ruler

The 86% in the September release is the share reporting visibility into where their data resides. So how did the other sectors do? The April release carries figures on the same subject. 54% of telecommunications respondents said it is extremely true that they have full visibility into where their data resides, while 30% of financial services respondents and 31% in the public sector reported the same. On being able to access all their data at any time, telecommunications ran 51%, financial services 24% and the public sector 16%.

Two conditions come attached to the April sentence and not to the energy and utilities one: full visibility, and answering extremely true. Only the strongest box on the answer scale is counted, in other words. The September sentence carries no such limit. To line the two figures up, a reader would have to know which boxes the 86% sums over, and the release does not supply that. Even so, the September release printed a sentence saying this sector is in better shape than the others surveyed.

There is one more place inside the same index where the rulers diverge. Cloudera's landing page for the telecom report states that 89% of telecom respondents claim complete visibility into their data ecosystems. The April release put the same sector, on the same subject, at 54%. When Cloudera set the sector results out again on its own blog on April 22, that value was still 54%, which makes a single typo hard to credit. Infrastructure performance hindering operations runs 60% in April and 90% on the telecom report page. Some items do agree across the two documents, the 33% for fully governed data among them, so calling them wholly different surveys is not available either. And within that single landing page, 89% claim complete visibility into their data ecosystems, while 60% admit their AI initiatives are hindered by an inability to access 100% of the data they need.

Item April global release Telecom report page
Full visibility into where data resides 54% (extremely true) 89%
Infrastructure performance hinders operations 60% 90%
All data fully governed 33% 33%

Values from two Cloudera documents on a single sector, telecommunications. The figures come from a survey of the same name, and two of the items are far apart. Neither document explains the difference.

The four sector cuts, set side by side, make the divergence of rulers plainer. On visibility into where data resides the figures are 89% on the telecom landing page, 87% in the healthcare announcement, 86% in the energy and utilities announcement and 82% in the manufacturing announcement. All four sit between 82% and 89%. The April release, on the same subject, gives telecommunications 54%, the public sector 31% and financial services 30%, and those come with the condition of answering extremely true about full visibility. Telecommunications appears on both sides, at 54% and at 89%. So the 86% does not belong on the same line as the April values, which run from the 30s to the 50s, and it is one of the 80s values that the sector announcements share among themselves.

The sentence the September release wrote without a number has the same problem. It is the infrastructure disruption sentence seen earlier, and the value that would serve as the comparison differs from document to document. The April release records 73% of all respondents as reporting that performance constraints have hindered operational initiatives, and the telecommunications figure in that same document is 60%. The telecom landing page says 90%, the healthcare announcement 28%. The question itself drifts between consistently hinders and has hindered. So which of these the September sentence took as its baseline is not clear.

To put it together, then. Three things have to match before sector figures can be set against one another: the wording of the question, the range of answer boxes summed, and the sample. For the sector cuts of this survey, no document confirms any of the three. So the 86% cannot serve as evidence that energy and utilities is better off than other sectors, and the 65% cannot be read as one notch below healthcare's 66%.

The figure in this survey that really does separate the sectors lies on the failure side instead. In energy and utilities the leading reason AI and analytics investments failed to deliver expected returns was cost overruns, at 25%. Across the full sample the leader was data quality at 22%, with cost overruns at 16%, below it. In healthcare, manufacturing and financial services, poor integration into existing workflows stood at the front at 20%. Because this is a comparison between sectors inside one questionnaire with identical wording, the contrast reads more easily than the 86% does.

Two roads lead out of that 25%. One reads it as a sector whose data foundation genuinely is in place, with money the remaining bottleneck. The other reads it as a data problem recorded under a cost line. When models run on data that is not ready, cleanup and rework, duplicate storage and retraining all cost more, and the overage lands in the books as a cost overrun rather than as a data quality problem. Which of the two holds cannot be settled from this survey, though the 84% confidence and the 18% full governance coming from the same sample put weight on the second.

4

What a Team Can Count Instead of Asking

From here the release is left behind, and the subject becomes what to measure when the same question is put to one's own organization. A vendor survey has to ask for a judgment, because a questionnaire is what it is. Inside an organization, records can be counted in place of judgments. The work of turning a yes into evidence usually starts with five counts.

  • Whether the bar word in our own report is all or nearly all. Until that is settled, every share that follows is open to interpretation.
  • The proportion of datasets where the governance policy is actually applied. Not the list of targets written into the policy document, but only the assets whose enforcement shows up in the policy engine and the access logs.
  • The median time from an access request to its approval, and the distribution of reasons behind the requests that were denied. The distance between an answer of yes, we can reach it and a record of an actual access shows up here.
  • The proportion of assets whose lineage runs unbroken back to the source. Where an asset sits and where it came from are different questions.
  • Whether another team, recalculating the same metric, arrives at the same value. A readiness score that does not reproduce is one person's impression.

All five can be pulled from audit logs and catalogs, and the value does not shift with whether the person pulling them views their own organization generously or harshly. That is where they part from self-reporting. An answer that the data is ready moves when the respondent changes; the number of datasets under an applied policy does not.

For anyone who has to keep reading vendor surveys, the questions come down to three. How many responses is this sector figure based on, how was the question worded, and how many boxes were summed. When even one of the three is absent from the document, that share cannot be used for comparison between sectors, and copying it into an internal report as a benchmark is risky.

5

Why Pebblous Is Watching This Survey

The question Pebblous has held on to for a long time is where a value came from and what it passed through to reach the shape it now has. The way the phrase data readiness is used at present sits exactly where that question was stepped around. In this survey readiness is not an observed state but a value manufactured out of answers. How that value was made ought to travel alongside it, and the release keeps only the percentage.

The moment this structure becomes a problem in practice is clear enough. An organization that gave its own readiness the benefit of the doubt looks outside its data when AI results fail to arrive. The budget ran over, so it writes cost overrun; the people could not keep up, so it writes a training problem. The data itself has already been answered for as ready, and so it is missing from the list of candidates. Something like that may have happened in the place where cost overruns rose to the top in energy and utilities. This is the reading of this article, and no such causation is written into the survey documents.

It is easy to point at a vendor's interests in a vendor survey, and on its own that yields little. Cloudera sells governance and a hybrid data platform, so it has an interest in diagnosing a governance gap. The part of this survey that deserves a longer look runs the other way. A question design that lets respondents hand their own organization a good score works against the vendor, since the reason to buy shrinks when everyone answers that they are ready. So the 86% and the 65% in this survey look less like a product of sales logic and more like a trace left by the method of self-reporting. The question that remains here is what a self-administered measurement makes invisible.

For a team that has put a data readiness check in front of its AI projects, four things are worth one look.

  • Who answered our readiness score. If the team that manages the data answered for itself, that score is a self-assessment.
  • How many instances of the word nearly are inside that score. Items carrying nearly all, broadly or most are values to count again on the all bar.
  • Whether data appears on last quarter's list of reasons AI results fell short. If it does not, it is worth checking whether an answer of ready is what kept it off the list.
  • Whether an outside survey figure has ever been copied into an internal target. Whether that figure's question and sample were ever matched against our own question and sample belongs to the same look.

Thank you for reading this far. The shares and the methodology this article cited can be checked by anyone in the originals, the September 17 energy and utilities release and the April 14 global release. We would be glad to hear which records your own organization uses to back a statement that its data is ready.

R

References

Cloudera Official Announcements & Press Releases

Industry Campaign Pages & Sponsored Content