Dataset opportunity
Tolid — LLM-geannoteerde corpora: crypto-nieuws, biomedische patenten, beautyvideo
Een dataonderneming die de corpora verkoopt die het heeft gebouwd — en nog steeds de pijplijn exploiteert die ze heeft gebouwd.
Score
84
Score (0–100) blends weighted dimensions — dataset rarity, training value, buyer demand, evidence strength and right-to-license. 70+ is deal-ready. See the scored dimensions below for the breakdown.Confidence
60%
Action
Licentie
The recommended deal structure for this dataset: Acquire (full buyout), License (paid usage rights), Data Sharing Agreement (controlled access, no transfer of ownership), Partnership (co-development) or Annotation Program (labeling). Chosen from data ownership, licensing complexity and accessibility.Market size (indicative estimate)
Geannoteerde trainingsdata waarvan de labelkwaliteit GEMETEN is, niet geclaimd — een claim die bijna geen enkel concurrerend corpus kan waarmaken. Verkocht met de gebreken verklaard, en met de pijplijn die het heeft geproduceerd beschikbaar als dienst.
Lineage
How this lead was derived
The signal-first chain, end to end: recent external signals → qualified niche → resolved data-holder → site verification → scored opportunity. Every lead is explainable.
Profile
Dataset profile
Type
LLM-geannoteerde corpora / kennisgrafen
Modality
tekst
Sector
AI-trainingsdata
Volume
2,47M docs · 723k patenten · 449 routines
Freshness
Archief — pijplijnen herstartbaar
Rarity
Hoog — de annotatie, niet de tekst
Accessibility
Onmiddellijk — BigQuery / Firestore spiegels
Legal
Annotaties, grafen en metadata zijn afgeleide werken geproduceerd door Tolid en zijn overdraagbaar. De bronpers-tekst is NIET (het behoort toe aan de uitgevers) — het is uitgesloten van elke levering. BIOPORTAL — gemeten en geregeld (2026-07-14): 707 ontologieën worden gebruikt, geen enkele. De 13.757 rijen (1,6% van de mappingtabel) afkomstig van restrictief gelicentieerde bronnen — SNOMED CT, MedDRA, RxNorm, OMIM, NDDF, Read Codes, ICD, CPT — zijn UITGESLOTEN van elke levering. Al het andere is open by construction (OBO Foundry eist CC-BY/CC0 van zijn leden; MeSH is gratis voor commercieel gebruik). CC-BY is een ATTRIBUTIEverplichting, geen copyleft: een ATTRIBUTIEKENNISGEVING wordt meegeleverd met de data. NOG OPEN: de Open Beauty Facts ODbL-status (beauty corpus) — een adversiële beoordeling weerlegde de 'Produced Work'-lezing, dus de beperking kan BINAIR blijken te zijn in plaats van een prijsvermindering. De vier biometrische inferentiekolommen (waargenomen geslacht, leeftijdsbereik, emotie, huidskleur) zijn uitgesloten van elke levering.
Buyer persona
Uitgevers van domein LLM's en financiële NLP-modellen · farma- en biotech-R&D (patentintelligentie) · cosmetica-groepen en beauty-tech · dataverkopers die hun eigen classificator willen trainen. NIET hedgefondsen of signaalverkopers voor het crypto-corpus — het draagt geen alpha, en we zeggen dit.
Tolid is een dataonderneming. Het verkoopt geen uitval van iemand anders door: het heeft de collectie- en LLM-annotatiepijplijnen gebouwd die deze corpora hebben geproduceerd, en het exploiteert ze nog steeds. Wat wordt aangeboden, zijn daarom twee dingen tegelijk — een voorraad geannoteerde data, en de machine die het heeft geproduceerd.
Drie corpora worden aangeboden.
1. Crypto-nieuws, geannoteerd (2.469.218 documenten, okt 2023 → sep 2025). Sentiment, onderwerp, samenvatting en een koop/houd/verkoop-tag, op genormaliseerde tickers. Lees dit vóór al het andere: het handelssignaal heeft GEEN gemeten voorspellende kracht. We hebben de backtest zelf uitgevoerd tegen reële prijzen, op alle 1.198.479 signalen: het verslaat een muntworp met 0,7 punten, en de marktneutrale alpha is nihil. De reden ligt in de data — perssentiment correleert met +0,60 met VORIGE rendementen en +0,07 met toekomstige. Het volgt de prijs; het loopt er niet op vooruit. We verkopen geen alpha, en dit corpus is niet voor een handelsdesk. Wat diezelfde +0,60 wel bewijst, is dat de annotatie getrouw is: het is een LLM-gelabeld financieel corpus waarvan de labelkwaliteit is gevalideerd tegen een externe grondwaarheid. Dat is wat u koopt — trainingsdata met een gemeten kwaliteit, geen signaal.
2. Biomedische patenten, gestructureerd als een kennisgraaf (723.149 documenten). Het meest waardevolle bezit van de drie, en degene met de grootste openstaande vraag: de licentievoorwaarden van de BioPortal-ontologieën waaraan het voldoet, worden herzien. We zeggen dit voordat u het vraagt.
3. Procedurele beautyvideo (449 routines, 7.656 getimede gebaren, 21.441 video's). Niet "videodata": een analytisch oppervlak over wat mensen daadwerkelijk doen, stap voor stap, met welke producten. Twee beperkingen vooraf verklaard — de Open Beauty Facts ODbL-kwestie is onopgelost en kan binair zijn in plaats van een korting, en de vier biometrische inferentiekolommen (waargenomen geslacht, leeftijd, emotie, huidskleur) zijn uitgesloten van elke levering. De procedurele graaf behoudt al zijn waarde zonder deze.
Wat we niet verkopen: de brontekst van het artikel. Het behoort toe aan de uitgevers, niet aan Tolid. Wat overdraagbaar is, is het afgeleide werk — de annotaties, de URL's, de metadata, de graaf.
Services
What the holder can also do for you
This dataset is not only a stock — its holder still operates the pipeline that produced it. Each capability below is backed by an observed fact, not a claim.
Human validation of the annotations, on demand
Tolid can run a human review pass over any corpus it has produced — full or sampled, to your own guidelines and quality bar.
Evidence — The entire human-in-the-loop layer already exists — nine moderation tables, an escalation queue, prompt structures — and has never been run on a single document (0 of 4,841,938). The machine is built and wired; it has simply never been switched on.
A corpus built to your theme, from scratch
Give a domain and a taxonomy; Tolid runs collection, LLM annotation and graph construction end-to-end — the same pipeline that produced these corpora.
Evidence — The pipeline is theme-driven, and there is proof: a seventh theme ("Geopolitical Tensions and Alliances") is fully configured — prompt and JSON template written — with no collection and no deployed service. A theme conceived, written, never launched. That is what "give us a subject" looks like in this codebase.
Entity resolution and normalisation
Mapping the messy surface forms of a domain onto canonical entities — the unglamorous work that decides whether a corpus is joinable with your own data.
Evidence — 21,359 alias → canonical-ticker mappings were built for the crypto corpus; 19,080 semantic mappings for the beauty ontology. This is the domain's price of entry, already paid.
Cleaning and extending an existing signal
Re-parsing, de-duplicating and extending a field that was produced once and never curated.
Evidence — Declared openly: the buy/hold/sell field covers only 33.4% of the crypto rows (1,198,479 of 3,588,282 — the other 66% read `Not mentioned`) and carries parsing leaks (`buy (implied)`, `hold/sell`). Tolid can clean and extend it. We would rather sell you the fix than hide the defect.
Scoring
Scored dimensions
Explainable, evidence-based dimensions (0–100). The radar shows the investment axes.
-
See dimension details ↓- Dataset Specificity88
Annotatiediepte, niet ruwe tekst: sentiment, onderwerp, samenvatting, triples — op 2,5 miljoen documenten, plus een kennisgraaf van 723.000 patenten.
How sharply the data targets a specific, hard-to-substitute domain or task. Niche, well-defined data scores higher than generic. - Dataset Rarity88
De tekst is overvloedig; de annotatie op deze diepte en geschiedenis is dat niet. Motormedian over drie runs: 85–90.
How scarce and proprietary the data is. Unique domain data scores high; openly available data lowers it. - Dataset Volume90
2.469.218 geannoteerde crypto-documenten · 723.149 patentdocumenten · 449 procedurele beauty-routines met 7.656 getimede gebaren.
Apparent scale of the data, inferred from the number of evidence hits and any explicit volume mentions. - Training Value85
Gebouwd voor modeltraining, niet voor rapportage — en de labelkwaliteit is gemeten, niet geclaimd (zie de crypto-validatiestudie).
How useful the data is for the target AI use-case — its fit for model training or fine-tuning. - Buyer Demand87
Motormedian 85–90 zodra de vraag wordt gewogen op basis van de solvabiliteit van de koper in plaats van op het aantal namen.
How strongly AI builders and companies are likely to want this data, based on market signals. - Evidence Strength45
Opzettelijk laag: het prijsbereik komt van onze waarderingsmotor, en GEEN externe marktvergelijkbare anker het nog. Twee diepgaande onderzoeken zijn in afwachting. We tonen liever een zwakke score dan een zelfverzekerd getal dat we niet kunnen verdedigen.
How solid the proof is that the company holds this data — diversity of evidence types and number of hits. - Data Orientation95
Een dataonderneming: de pijplijnen, de graaf en de annotatielaag zijn het product, geen bijproduct.
How actively the company invests in data, measured by its data-appetite signals (hires, products, APIs…). - Dormant Data Surplus92
Geproduceerd, opgeslagen — en nooit gemonetiseerd. De moderatielaag is nooit actief geweest; het crypto-corpus is bevroren sinds 2025-09-20.
Volume and value of proprietary data this company holds BEYOND what it already monetises — the dormant surplus we can unlock. A company can sell some insights AND still sit on a far larger dormant asset.
Evidence
Dataset evidence & lineage
What the typed evidence proves the company holds — reframed for clarity and set against the market.
-
Marketplace
Dataset details
Geographic coverage
Global (multilingual sources)
Time range
2023-10 → 2025-09 (crypto) · patents and beauty: see each dataset
Delivery
Secure extract download, or scoped access
Formats
parquet, csv, jsonl
License
Derived works (annotations, graphs, metadata) are transferable. Source press text is not.
Personal data
No PII
There is no public price grid for these corpora, and we will not manufacture one. Our own valuation engine returns wide, UNANCHORED ranges — dispersion up to ×7 on the crypto asset across three identical runs — and a dedicated deep-research study (July 2026, 103 agents, 86 claims → 25 adversarially verified) confirmed WHY: for an LLM-annotated news archive sold as annotations-only, no defensible market comparable exists. The published editor↔LLM deals disclose amounts but never volumes, so no per-document price can be derived; and the closest academic comparable (Financial PhraseBank) anchors label PEDIGREE, never a number. The gap is in the market, not in our research. So we price ON REQUEST, against two things we can defend: a measured production cost (from real cloud billing) and a label quality validated against external ground truth. The biomedical-patents asset — the strongest of the three — is under its own valuation study; its price will be shown only once anchored on real dataset transactions, never on a SaaS-subscription proxy. You do not buy the bundle: each dataset is priced on its own.
Detailed schema & sample available on access request.
Want this data?
Request access — we broker a secure deal room. Operator-reviewed, no automatic sharing.
This listing was generated automatically from public signals. It is not verified, and we are not affiliated with this company.
Coverage
Scanned sources
Deliverable
Premium dataset report
A teaser is generated for each dataset opportunity. The full report is available on unlock.
From the marketplace
Explore live data opportunities
Accordiaglobal — Dataset Gelegenheid voor Inspectierapporten
View opportunity →gezondheidszorgFusixbiotech — Gelegenheid voor dataset industriële operaties
View opportunity →mobiliteitTuenvioya — Mogelijkheid voor dataset openbare aanbestedingen
View opportunity →Data Academy
Learn before you deal
- Hoe een datatransactie verloopt3 min read
- Wat u mag verkopen3 min read
- 5 fouten die kopers afschrikken3 min read