Startup Spotlight: Petrarch is turning industrial records into frontier AI training data

Petrarch says it has transacted construction plans, autonomous-driving recordings and codebases; buyers, dataset volumes and transaction values remain undisclosed.

By · Published

Primary source: Y Combinator

Why it matters

Petrarch is betting that access to useful industrial data depends on the work between finding a file and making it usable: sourcing, rights review, privacy handling and technical preparation.

A close-up shot of an industrial record book being scanned by an advanced optical scanner, with data visualizations displayed on a nearby screen.

Petrarch wants to turn the records that companies use to run factories, projects and logistics operations into training material for AI systems. The bet draws on three founders whose previous work spans training-data quality, privacy tooling and computer vision: Ian Lee, Samuel Hahn and Sudhish Swain. In its Y Combinator profile, the San Francisco startup says it sources internal industrial data, including codebases, project files and payment records, and prepares it for model training.

Petrarch's early examples reach beyond manufacturing. The YC launch description says the founders have transacted construction plans, autonomous-driving recordings and codebases. That list is a company-reported account of activity, rather than a measure of scale: each example shows a different kind of information, but the profile does not describe the buyers, dataset sizes or transaction values. Petrarch's promise depends on turning one-off access to obscure records into a dependable supply of material labs can use.

Three founders, three parts of the problem

Lee's background puts data quality near the center of Petrarch's pitch. YC says he was a fellow at Meridian AI, working on training-data and go-to-market infrastructure for an AI-native spreadsheet product in financial services. The profile also identifies him as a technical adviser at Scale AI, where he worked on quality filtering and annotation for agentic assistants and reasoning models. He studied applied mathematics and psychology at Harvard, according to YC.

Hahn brings a different piece of the work: getting sensitive data into a form a buyer can consider using. YC says he led go-to-market at Magier AI, a Techstars 2024 company building agentic products for personally identifiable information redaction and data compliance. He also worked as a venture analyst at Mainstreet. Those experiences sit close to the commercial and handling questions in Petrarch's model: a supplier must trust the intermediary, and a buyer must believe the data has been prepared for its intended use.

Swain's prior work is more directly tied to physical operations. YC says he led software engineering at Clean Sweep Group, building computer-vision software to identify medical equipment deployed in California hospitals. He also researched spatial-transcriptomics cancer datasets at Brigham and Women's Hospital. That background links machine perception with the messy, specialized information produced in real workplaces - the kind of source material that is harder to collect than public text.

YC's launch copy says all three founders left Harvard before completing their studies. Their biographies make the founding team unusually aligned with the mechanics of data supply: Lee worked on filtering and annotation, Hahn on privacy-focused product sales, and Swain on applied vision and data research. The team's experience helps explain why Petrarch is building around the preparation and transfer of data rather than pitching another general-purpose model; it does not establish demand.

From files to usable AI tasks

The Petrarch website describes a pipeline that anonymizes and analyzes operational data, then restructures it into authentic tasks and real-world environments. Its examples of source systems include tools such as Notion, GitHub, Confluence, Grafana, Snowflake and SAP. The ambition is broader than handing a lab a folder of documents: records need to be organized so they can support training or evaluation that reflects how work actually happens.

That conversion step is the harder product problem. A project file can contain a useful sequence of decisions, but it may also include names, customer details, licensed material or confidential business information. A codebase can reveal how software was built while carrying its own history of ownership and third-party components. A payment record can encode vendor relationships and operating patterns. Petrarch says it de-identifies and prepares material for model training; the practical value of that promise depends on preserving useful context while separating it from information a supplier cannot authorize for reuse.

The sourcing model combines at least two channels. The YC profile highlights records obtained through bankruptcy courts, while Petrarch's homepage describes long-term partnerships with operating businesses. Petrarch also says it targets legacy businesses, startups and companies going through bankruptcy or pivots. That range gives Petrarch more potential sources than a marketplace limited to court proceedings, but it also means the company must handle different permissions and expectations depending on how each dataset is obtained.

Bankruptcy is an especially distinctive source because a failed or restructuring business may leave behind useful records that were never designed for AI training. But access to a file is not the same thing as a clear right to process and license every piece of information inside it. Petrarch's value proposition therefore rests on more than discovery. Its process has to establish what the records contain, what can be retained or transformed, and what a buyer is permitted to do with the resulting dataset. That is a central part of the product, not administrative work around the edges.

A market that still needs proof of repeatability

Petrarch is entering a data business where the hardest assets are often hidden inside institutions that did not build their systems for external reuse. Publicly available material and manually prepared examples remain available to AI developers, but proprietary operating records can capture procedures and edge cases that are difficult to reproduce from open sources alone. Petrarch's stated focus is on that less accessible layer, starting with manufacturing and industrial companies.

Its reported transactions offer an initial indication of range: construction plans, autonomous-driving recordings and codebases are different formats with different potential users. Petrarch says it sources, cleans and restructures material before reselling it to applied AI startups and labs. That model leaves a demanding operational task between supplier and buyer: converting inconsistent files into usable datasets while maintaining an understandable account of where they came from and how they may be used.

A YC founding-engineer listing gives a glimpse of how the team intends to build the service. It describes Lindo as software deployed into customer cloud environments and asks the hire to work across data ingestion, processing and customer-facing applications. The post also lists experience with enterprise cloud deployments, identity systems and high-volume data pipelines. Those requirements point toward customer-specific infrastructure work, rather than a self-serve catalog where a buyer simply downloads a dataset.

That approach could help with trust and security: customers may be more comfortable when sensitive records can be processed within an environment they control. It also adds deployment work for each customer and places weight on integration, access controls and repeatable data handling. The hiring post describes the intended engineering remit; it does not establish how many deployments are operating or how much of the pipeline is automated.

The public materials also place Petrarch between two constituencies with different incentives. Businesses may have records that are costly to organize and difficult to monetize, while AI developers want specialized material without building supplier relationships one at a time. Petrarch's intermediary role is to make that exchange possible and usable. Its advantage, if it develops one, will come from reliable sourcing and preparation rather than exclusive access to any single category of file.

What the founders are building

YC lists Petrarch as an active Summer 2026 company with three founders. Its public materials say the startup has completed transactions across three data categories and is building tools to prepare internal company records for AI development. The evidence supports an early business taking shape around a specific bottleneck; it does not yet establish the scale or repeatability of that business.

That makes the founders' experience more than a credential list. Lee has worked on the quality layer, Hahn on selling and handling data-sensitive products, and Swain on software built for physical-world environments. Petrarch is attempting to bring those capabilities together around a category of information that businesses have accumulated for years but rarely package for outside use.

The commercial test is whether suppliers will permit that reuse and whether AI developers will pay for records that have been cleaned, contextualized and cleared for a defined purpose. Petrarch's examples suggest the founders have begun working across several kinds of data. The next measure is whether those individual transactions turn into a repeatable pipeline - one that can deliver usable material without asking each buyer and supplier to solve the same sourcing and preparation problems from scratch.

Reader comments

Conversation for this story loads after sign-in.