OpenAI launches data agent to turn enterprise data into analytics and dashboards with a single command, but accuracy benchmarks remain undisclosed

Deep News
Sep 11

OpenAI has taken another step into the enterprise data analytics market with a new offering aimed at simplifying how businesses interact with their information.

On September 10, local time, OpenAI announced the launch of a new "Data agent" feature within ChatGPT Work. This tool allows users to skip writing SQL code or manually consolidating multiple data sources. Instead, they can pose questions in natural language, and the agent will connect to authorized corporate data and documents, investigate what has changed within the data, generate analysis and interactive dashboards, and execute follow-up actions after receiving user approval.

However, the product launch leaves a key question unanswered: OpenAI has not disclosed any public accuracy or retrieval accuracy benchmarks for external customers. This means enterprises can see which data sources the agent can connect to and what tasks it can accomplish, but they currently lack an independent, quantifiable public metric to determine just how accurate its answers truly are.

A single question can interrogate 600PB of data, turning data analysis into a conversation.

OpenAI has stated that the goal of the data agent is to let employees directly "ask questions of their data." For instance, users can inquire why a certain business metric has changed, request the system to identify contributing factors, compare different time periods or customer segments, and then generate reports needed by management. The agent will investigate relevant data on its own, rather than simply returning a piece of SQL code or a single number.

It can also transform analysis results directly into interactive dashboards, including charts, metrics, and editable visualizations, while supporting team sharing, refreshing, and follow-up questions within the team.

Prior to officially announcing the new agent, OpenAI's internal data agent had already been operating on a massive scale. According to reports from VentureBeat, OpenAI's internal data agent has served more than 3,500 users, covering over 600PB of data and roughly 70,000 datasets.

This helps explain why OpenAI emphasizes that this product was not developed from scratch, but rather represents a further productization of the data analysis workflows already in use internally.

From database to dashboard: bridging enterprise data, files, and BI tools.

Compared with traditional "natural language query database" tools, OpenAI places greater emphasis on how the data agent integrates with a company's existing toolchain. Currently supported data platforms include Amazon Redshift, Google BigQuery, ClickHouse, Databricks, MongoDB, Snowflake, and Datadog. It can also incorporate files and documents from Google Drive and SharePoint into its analysis.

Once analysis is complete, the data agent can collaborate with dashboard and BI tools such as Omni, Oracle BI, Power BI, Sigma, Tableau, and ThoughtSpot.

OpenAI states that the data agent can also leverage business terminology, metric definitions, custom calculation methods, and data relationships already established within an enterprise. In other words, it seeks to understand not just "what is in this table," but how the company internally defines business metrics such as revenue, customers, and orders.

Additionally, administrators can pre-determine which data connections and roles have access, and data queries continue to follow the enterprise's existing table-level, row-level, and column-level permissions.

Going a step further, the data agent is not limited to just "providing answers." Based on analysis results, it can suggest who should be contacted next, which teams need to be involved, and can share results or execute user-approved actions through connected tools.

First use case is internal: over 3,500 people at OpenAI already rely on it for data queries.

OpenAI says the data agent originated from the company's own data usage needs. VentureBeat quoted OpenAI's enterprise technology leader, Arpan Shah, as saying that about a year ago, only a very small number of data questions at OpenAI could be resolved end-to-end without human intervention. Today, employees can handle a large number of data questions on their own.

The internal version has served more than 3,500 users, covering approximately 70,000 datasets and over 600PB of data. OpenAI also indicates that currently, nearly all product team members and more than two-thirds of the GTM (go-to-market and sales) team are using the data agent.

This also marks a significant shift in OpenAI's product positioning: it aims to turn work that previously required data analysts and engineers into tasks that ordinary employees can complete through natural language.

The biggest question mark: it can answer questions, but how accurate is it really?

For enterprise-level data products, however, "can it do it" is only the first hurdle; "how accurately does it do it" is the more critical issue.

OpenAI did not release a public accuracy or retrieval accuracy benchmark for external customers during this launch. VentureBeat reported that OpenAI did conduct internal comparison tests, including comparing the data agent's results against the company's existing data tools, but OpenAI did not disclose specific scores, nor did it provide a number that external customers could use for horizontal comparison.

This does not necessarily mean the data agent has low accuracy. Rather, it means enterprise customers currently lack an open, unified quantitative standard to judge its performance in complex corporate data environments.

This point deserves particular attention because enterprise data is often far more complex than standard Q&A: answers may be scattered across multiple data tables, BI dashboards, documents, and even Slack messages, and different teams may define the same metric in different ways.

At the same time, competitors are actively using benchmarks to prove their own data intelligence capabilities. For example, Databricks recently published research on its Adaptive Instructed-Retriever, claiming that its system's answer quality on specific enterprise retrieval tasks can reach the level of Claude Sonnet 5, GPT-5.6 Luna, and DeepSeek-V4-Flash, with an average response time of approximately 5.8 seconds. It should be noted that these results come from Databricks' own testing and have not yet been independently verified.

Consequently, as enterprises begin entrusting more data queries to agents, whoever can prove "accuracy" with public, reproducible metrics may be just as important as whoever can connect to more data sources.

OpenAI is pushing data intelligence into the enterprise, aiming to reduce dependence on analysts for every query.

From OpenAI's product design perspective, the ultimate goal is not simply to launch a "stronger SQL assistant," but to make the data agent a new entry point between enterprise employees and their data.

Sales professionals can directly analyze customer and sales data, marketing teams can track campaign metrics, and management can ask the agent to investigate business changes and generate dashboards, rather than waiting each time for a data analyst to prepare a report.

This also positions the data agent as a further step for ChatGPT Work to penetrate deeper into enterprise workflows: moving from answering questions to investigating data, creating analyses, generating visualizations, and then executing follow-up actions.

On the other side, enterprise data infrastructure vendors such as Databricks and Snowflake are also accelerating their push toward agents. The future competition in enterprise data intelligence may no longer be just about "who has the better large language model," but about who can understand enterprise data more accurately, adhere to permission systems, and truly translate analysis results into business actions.

For OpenAI, this launch addresses the problem of "letting everyone ask questions of data." As for whether it can further prove "letting everyone trust the answers with confidence," a public accuracy benchmark remains the next test it still needs to face.

Disclaimer: Investing carries risk. This is not financial advice. The above content should not be regarded as an offer, recommendation, or solicitation on acquiring or disposing of any financial products, any associated discussions, comments, or posts by author or other users should not be considered as such either. It is solely for general information purpose only, which does not consider your own investment objectives, financial situations or needs. TTM assumes no responsibility or warranty for the accuracy and completeness of the information, investors should do their own research and may seek professional advice before investing.

Most Discussed

  1. 1
     
     
     
     
  2. 2
     
     
     
     
  3. 3
     
     
     
     
  4. 4
     
     
     
     
  5. 5
     
     
     
     
  6. 6
     
     
     
     
  7. 7
     
     
     
     
  8. 8
     
     
     
     
  9. 9
     
     
     
     
  10. 10