According to sources familiar with the matter, SpaceX (SPCX.US) has held internal discussions about purchasing customer and operational data from struggling or defunct startups, aiming to secure high-quality datasets at relatively low costs to enhance the performance of its artificial intelligence models. The talks are currently taking place primarily within SpaceXAI, the company's AI division, and remain in an informal stage, with no guarantee that any deal will ultimately be reached. SpaceX did not respond to requests for comment.
This approach bears a striking resemblance to a strategy previously employed by Google. Earlier this year, after U.S. budget airline Spirit Airlines ceased operations, Google proposed acquiring its commercial data for $10 million to use in AI training. However, that potential transaction also raised data privacy concerns, including objections from some former Spirit Airlines flight attendants.
Where the search for new data stands
SpaceXAI, formerly known as xAI, is intensifying its competition with AI rivals such as Anthropic and OpenAI as Musk's space, satellite communications, and artificial intelligence ventures become more integrated, while the company simultaneously seeks to attract a broader corporate client base. Like other AI firms, one of SpaceXAI's key challenges is sourcing high-quality training data that can improve model performance across diverse tasks. According to insiders, the company is particularly interested in enterprise operational data and customer information, with plans to feed these external datasets into models like Grok.
If implemented, this would mark a notable shift in SpaceXAI's data strategy. Historically, the company has relied primarily on data from Musk's social media platform X, supplemented by a team of internal specialists serving as "AI tutors" who help train and refine AI models. These AI tutors assist engineers in improving model capabilities in specialized domains such as finance, science, and even humor. The current consideration of external data sources reflects a growing recognition that, as AI model capabilities advance, relying solely on internal data may no longer fully satisfy training and optimization requirements. For leading AI enterprises, legally acquiring data that is professional, authentic, and structured is becoming an increasingly vital resource for enhancing model performance.
Restructuring the AI training team
Concurrently, the data team responsible for model training within SpaceXAI has undergone a series of personnel and strategic adjustments in recent months. In June, SpaceX temporarily halted hiring for AI tutors responsible for training Grok, and subsequently reshuffled team leadership. Jack Garabedian, a longtime veteran of Starlink, SpaceX's satellite internet division, recently took over the team, replacing younger engineer Diego Pasini. Sources indicate that since assuming his role, Garabedian has been working to improve internal operations, including establishing clearer workflows, meeting protocols, and AI training data targets.
An internal communication from SpaceXAI reviewed by media outlets revealed that the data team's July progress included advancing development of a new AI model and a coding agent. The communication emphasized that without the team's daily foundational work in annotation, preference alignment, and validation, these models would struggle to evolve into mature products. This underscores that data processing and human feedback remain critical components of AI model training, alongside computing power and model architecture.
Balancing external procurement with internal data streams
While SpaceXAI is exploring additional external data sources, Musk's vast business empire itself remains a significant wellspring of data for Grok, including information generated by SpaceX employees. During a recent internal SpaceX meeting, Musk stated that the company plans to train Grok using all internal SpaceX information, directly telling employees, "It will also train on your data." This statement suggests that SpaceXAI's future data strategy may adopt a dual-track approach, combining internal and external sources: continuing to leverage data generated by Musk-owned enterprises and platforms like SpaceX and X, while simultaneously expanding training data breadth and professionalism through external dataset acquisitions.
It is worth noting that while acquiring data from distressed or defunct companies may offer a relatively low-cost source of information, the use of customer data also raises issues related to privacy, authorization, and permissible data usage scope. The controversy surrounding Google's earlier bid for Spirit Airlines data has already highlighted the potential compliance and privacy risks inherent in such transactions. As OpenAI, Anthropic, and SpaceXAI compete over model performance and enterprise clients, high-quality, specialized data is emerging as yet another critical battleground, following chips, computing power, and electricity in the AI race. SpaceX's current consideration of purchasing operational and customer data from struggling startups also signals that the competition among leading AI firms for training data is extending further into the realm of private corporate data.