Dataset opportunity

Tolid — LLM-geannoteerde corpora: crypto-nieuws, biomedische patenten, beautyvideo

Een dataonderneming die de corpora verkoopt die het heeft gebouwd — en nog steeds de pijplijn exploiteert die ze heeft gebouwd.

LLM-geannoteerde corpora / kennisgrafentekstDomein LLM-training, kennisgrafen, sentimentmodellen🌍 Francetolid.io16 jul 2026

Confidence

60%

Market size (indicative estimate)

Geannoteerde trainingsdata waarvan de labelkwaliteit GEMETEN is, niet geclaimd — een claim die bijna geen enkel concurrerend corpus kan waarmaken. Verkocht met de gebreken verklaard, en met de pijplijn die het heeft geproduceerd beschikbaar als dienst.

Lineage

How this lead was derived

The signal-first chain, end to end: recent external signals → qualified niche → resolved data-holder → site verification → scored opportunity. Every lead is explainable.

Profile

Dataset profile

Type

LLM-geannoteerde corpora / kennisgrafen

Modality

tekst

Sector

AI-trainingsdata

Volume

2,47M docs · 723k patenten · 449 routines

Freshness

Archief — pijplijnen herstartbaar

Rarity

Hoog — de annotatie, niet de tekst

Accessibility

Onmiddellijk — BigQuery / Firestore spiegels

Legal

Annotaties, grafen en metadata zijn afgeleide werken geproduceerd door Tolid en zijn overdraagbaar. De bronpers-tekst is NIET (het behoort toe aan de uitgevers) — het is uitgesloten van elke levering. BIOPORTAL — gemeten en geregeld (2026-07-14): 707 ontologieën worden gebruikt, geen enkele. De 13.757 rijen (1,6% van de mappingtabel) afkomstig van restrictief gelicentieerde bronnen — SNOMED CT, MedDRA, RxNorm, OMIM, NDDF, Read Codes, ICD, CPT — zijn UITGESLOTEN van elke levering. Al het andere is open by construction (OBO Foundry eist CC-BY/CC0 van zijn leden; MeSH is gratis voor commercieel gebruik). CC-BY is een ATTRIBUTIEverplichting, geen copyleft: een ATTRIBUTIEKENNISGEVING wordt meegeleverd met de data. NOG OPEN: de Open Beauty Facts ODbL-status (beauty corpus) — een adversiële beoordeling weerlegde de 'Produced Work'-lezing, dus de beperking kan BINAIR blijken te zijn in plaats van een prijsvermindering. De vier biometrische inferentiekolommen (waargenomen geslacht, leeftijdsbereik, emotie, huidskleur) zijn uitgesloten van elke levering.

Buyer persona

Uitgevers van domein LLM's en financiële NLP-modellen · farma- en biotech-R&D (patentintelligentie) · cosmetica-groepen en beauty-tech · dataverkopers die hun eigen classificator willen trainen. NIET hedgefondsen of signaalverkopers voor het crypto-corpus — het draagt geen alpha, en we zeggen dit.

Tolid is een dataonderneming. Het verkoopt geen uitval van iemand anders door: het heeft de collectie- en LLM-annotatiepijplijnen gebouwd die deze corpora hebben geproduceerd, en het exploiteert ze nog steeds. Wat wordt aangeboden, zijn daarom twee dingen tegelijk — een voorraad geannoteerde data, en de machine die het heeft geproduceerd.

Drie corpora worden aangeboden.

1. Crypto-nieuws, geannoteerd (2.469.218 documenten, okt 2023 → sep 2025). Sentiment, onderwerp, samenvatting en een koop/houd/verkoop-tag, op genormaliseerde tickers. Lees dit vóór al het andere: het handelssignaal heeft GEEN gemeten voorspellende kracht. We hebben de backtest zelf uitgevoerd tegen reële prijzen, op alle 1.198.479 signalen: het verslaat een muntworp met 0,7 punten, en de marktneutrale alpha is nihil. De reden ligt in de data — perssentiment correleert met +0,60 met VORIGE rendementen en +0,07 met toekomstige. Het volgt de prijs; het loopt er niet op vooruit. We verkopen geen alpha, en dit corpus is niet voor een handelsdesk. Wat diezelfde +0,60 wel bewijst, is dat de annotatie getrouw is: het is een LLM-gelabeld financieel corpus waarvan de labelkwaliteit is gevalideerd tegen een externe grondwaarheid. Dat is wat u koopt — trainingsdata met een gemeten kwaliteit, geen signaal.

2. Biomedische patenten, gestructureerd als een kennisgraaf (723.149 documenten). Het meest waardevolle bezit van de drie, en degene met de grootste openstaande vraag: de licentievoorwaarden van de BioPortal-ontologieën waaraan het voldoet, worden herzien. We zeggen dit voordat u het vraagt.

3. Procedurele beautyvideo (449 routines, 7.656 getimede gebaren, 21.441 video's). Niet "videodata": een analytisch oppervlak over wat mensen daadwerkelijk doen, stap voor stap, met welke producten. Twee beperkingen vooraf verklaard — de Open Beauty Facts ODbL-kwestie is onopgelost en kan binair zijn in plaats van een korting, en de vier biometrische inferentiekolommen (waargenomen geslacht, leeftijd, emotie, huidskleur) zijn uitgesloten van elke levering. De procedurele graaf behoudt al zijn waarde zonder deze.

Wat we niet verkopen: de brontekst van het artikel. Het behoort toe aan de uitgevers, niet aan Tolid. Wat overdraagbaar is, is het afgeleide werk — de annotaties, de URL's, de metadata, de graaf.

Services

What the holder can also do for you

This dataset is not only a stock — its holder still operates the pipeline that produced it. Each capability below is backed by an observed fact, not a claim.

Human validation of the annotations, on demand

Tolid can run a human review pass over any corpus it has produced — full or sampled, to your own guidelines and quality bar.

EvidenceThe entire human-in-the-loop layer already exists — nine moderation tables, an escalation queue, prompt structures — and has never been run on a single document (0 of 4,841,938). The machine is built and wired; it has simply never been switched on.

A corpus built to your theme, from scratch

Give a domain and a taxonomy; Tolid runs collection, LLM annotation and graph construction end-to-end — the same pipeline that produced these corpora.

EvidenceThe pipeline is theme-driven, and there is proof: a seventh theme ("Geopolitical Tensions and Alliances") is fully configured — prompt and JSON template written — with no collection and no deployed service. A theme conceived, written, never launched. That is what "give us a subject" looks like in this codebase.

Entity resolution and normalisation

Mapping the messy surface forms of a domain onto canonical entities — the unglamorous work that decides whether a corpus is joinable with your own data.

Evidence21,359 alias → canonical-ticker mappings were built for the crypto corpus; 19,080 semantic mappings for the beauty ontology. This is the domain's price of entry, already paid.

Cleaning and extending an existing signal

Re-parsing, de-duplicating and extending a field that was produced once and never curated.

EvidenceDeclared openly: the buy/hold/sell field covers only 33.4% of the crypto rows (1,198,479 of 3,588,282 — the other 66% read `Not mentioned`) and carries parsing leaks (`buy (implied)`, `hold/sell`). Tolid can clean and extend it. We would rather sell you the fix than hide the defect.

Scoring

Scored dimensions

Explainable, evidence-based dimensions (0–100). The radar shows the investment axes.

-

See dimension details
SpecificityRarityVolumeTraining ValueBuyer DemandEvidence StrengthData Orientation

Evidence

Dataset evidence & lineage

What the typed evidence proves the company holds — reframed for clarity and set against the market.

-

Marketplace

Dataset details

Geographic coverage

Global (multilingual sources)

Time range

2023-10 → 2025-09 (crypto) · patents and beauty: see each dataset

Delivery

Secure extract download, or scoped access

Formats

parquet, csv, jsonl

License

Derived works (annotations, graphs, metadata) are transferable. Source press text is not.

Personal data

No PII

· licence· final price on request

There is no public price grid for these corpora, and we will not manufacture one. Our own valuation engine returns wide, UNANCHORED ranges — dispersion up to ×7 on the crypto asset across three identical runs — and a dedicated deep-research study (July 2026, 103 agents, 86 claims → 25 adversarially verified) confirmed WHY: for an LLM-annotated news archive sold as annotations-only, no defensible market comparable exists. The published editor↔LLM deals disclose amounts but never volumes, so no per-document price can be derived; and the closest academic comparable (Financial PhraseBank) anchors label PEDIGREE, never a number. The gap is in the market, not in our research. So we price ON REQUEST, against two things we can defend: a measured production cost (from real cloud billing) and a label quality validated against external ground truth. The biomedical-patents asset — the strongest of the three — is under its own valuation study; its price will be shown only once anchored on real dataset transactions, never on a SaaS-subscription proxy. You do not buy the bundle: each dataset is priced on its own.

Detailed schema & sample available on access request.

Want this data?

Request access — we broker a secure deal room. Operator-reviewed, no automatic sharing.

Share this opportunity

This listing was generated automatically from public signals. It is not verified, and we are not affiliated with this company.

Coverage

Scanned sources

https://tolid.iodiscovered

Deliverable

Premium dataset report

A teaser is generated for each dataset opportunity. The full report is available on unlock.

Teaser is public · premium is locked behind access.

From the marketplace

Explore live data opportunities

Browse datasets by sector & use-case