Dataset opportunity

Tolid — Corpora annotati da LLM: notizie crypto, brevetti biomedici, video di bellezza

Un'azienda di dati che vende i corpora che ha creato — e continua a gestire la pipeline che li ha generati.

Corpora annotati da LLM / knowledge graphstestoAddestramento LLM di dominio, knowledge graphs, modelli di sentiment🌍 Francetolid.io16 lug 2026

Confidence

60%

Market size (indicative estimate)

Dati di addestramento annotati la cui qualità delle etichette è MISURATA, non affermata — un'affermazione che quasi nessun corpus concorrente può fare. Venduto con i suoi difetti dichiarati, e con la pipeline che lo ha prodotto disponibile come servizio.

Lineage

How this lead was derived

The signal-first chain, end to end: recent external signals → qualified niche → resolved data-holder → site verification → scored opportunity. Every lead is explainable.

Profile

Dataset profile

Type

Corpora annotati da LLM / knowledge graphs

Modality

testo

Sector

Dati di addestramento AI

Volume

2,47M doc · 723k brevetti · 449 routine

Freshness

Archivio — pipeline riavviabili

Rarity

Alta — l'annotazione, non il testo

Accessibility

Immediata — mirror BigQuery / Firestore

Legal

Annotazioni, grafi e metadati sono lavori derivati prodotti da Tolid e sono trasferibili. Il testo della stampa sorgente NON LO È (appartiene agli editori) — è escluso da ogni consegna. BIOPORTAL — misurato e risolto (2026-07-14): vengono utilizzate 707 ontologie, non una. Le 13.757 righe (1,6% della tabella di mappatura) provenienti da fonti con licenza restrittiva — SNOMED CT, MedDRA, RxNorm, OMIM, NDDF, Read Codes, ICD, CPT — sono ESCLUSE da ogni consegna. Tutto il resto è aperto per costruzione (OBO Foundry impone CC-BY/CC0 ai suoi membri; MeSH è gratuito per uso commerciale). CC-BY è un obbligo di ATTRIBUZIONE, non un copyleft: un avviso di attribuzione viene fornito con i dati. ANCORA APERTO: lo stato Open Beauty Facts ODbL (corpus bellezza) — la revisione avversaria ha confutato la lettura "Opera Prodotta", quindi il vincolo potrebbe rivelarsi BINARIO piuttosto che uno sconto sul prezzo. Le quattro colonne di inferenza biometrica (sesso percepito, fascia d'età, emozione, tono della pelle) sono escluse da qualsiasi consegna.

Buyer persona

Editori di LLM di dominio e modelli finanziari NLP · R&S farmaceutica e biotecnologica (intelligence brevettuale) · gruppi cosmetici e beauty-tech · fornitori di dati che desiderano addestrare il proprio classificatore. NON hedge fund o fornitori di segnali per il corpus crypto — non porta alpha, e lo diciamo.

Tolid è un'azienda di dati. Non rivende dati di altri: ha creato le pipeline di raccolta e annotazione LLM che hanno prodotto questi corpora, e continua a gestirle. Ciò che viene offerto sono quindi due cose contemporaneamente — un stock di dati annotati e la macchina che li ha prodotti.

Sono in offerta tre corpora.

1. Notizie crypto, annotate (2.469.218 documenti, Ott 2023 → Set 2025). Sentiment, argomento, riassunto e un tag buy/hold/sell, su ticker normalizzati. Leggere questo prima di tutto: il segnale di trading NON ha potere predittivo misurato. Abbiamo eseguito il backtest contro prezzi reali, su tutti gli 1.198.479 segnali: supera un lancio di moneta di 0,7 punti e il suo alpha market-neutral è nullo. La ragione è nei dati — il sentiment della stampa correla a +0,60 con i rendimenti PASSATI e +0,07 con quelli futuri. Segue il prezzo; non lo precede. Non vendiamo alpha, e questo corpus non è per una desk di trading. Ciò che lo stesso +0,60 dimostra è che l'annotazione è fedele: è un corpus finanziario etichettato da LLM la cui qualità delle etichette è stata validata rispetto a una ground truth esterna. Questo è ciò che state acquistando — dati di addestramento con una qualità misurata, non un segnale.

2. Brevetti biomedici, strutturati come knowledge graph (723.149 documenti). L'asset più prezioso dei tre, e quello che porta la più grande domanda aperta: i termini di licenza delle ontologie BioPortal a cui si allinea sono in fase di revisione. Lo diciamo prima che lo chiediate.

3. Video di bellezza procedurali (449 routine, 7.656 gesti timestamped, 21.441 video). Non "dati video": una superficie analitica su ciò che le persone fanno effettivamente, passo dopo passo, con quali prodotti. Due vincoli dichiarati in anticipo — la questione Open Beauty Facts ODbL è irrisolta e potrebbe essere binaria piuttosto che uno sconto, e le quattro colonne di inferenza biometrica (sesso percepito, età, emozione, tono della pelle) sono escluse da qualsiasi consegna. Il grafo procedurale mantiene tutto il suo valore senza di esse.

Cosa non vendiamo: il testo dell'articolo sorgente. Appartiene agli editori, non a Tolid. Ciò che è trasferibile è il lavoro derivato — le annotazioni, gli URL, i metadati, il grafo.

Services

What the holder can also do for you

This dataset is not only a stock — its holder still operates the pipeline that produced it. Each capability below is backed by an observed fact, not a claim.

Human validation of the annotations, on demand

Tolid can run a human review pass over any corpus it has produced — full or sampled, to your own guidelines and quality bar.

EvidenceThe entire human-in-the-loop layer already exists — nine moderation tables, an escalation queue, prompt structures — and has never been run on a single document (0 of 4,841,938). The machine is built and wired; it has simply never been switched on.

A corpus built to your theme, from scratch

Give a domain and a taxonomy; Tolid runs collection, LLM annotation and graph construction end-to-end — the same pipeline that produced these corpora.

EvidenceThe pipeline is theme-driven, and there is proof: a seventh theme ("Geopolitical Tensions and Alliances") is fully configured — prompt and JSON template written — with no collection and no deployed service. A theme conceived, written, never launched. That is what "give us a subject" looks like in this codebase.

Entity resolution and normalisation

Mapping the messy surface forms of a domain onto canonical entities — the unglamorous work that decides whether a corpus is joinable with your own data.

Evidence21,359 alias → canonical-ticker mappings were built for the crypto corpus; 19,080 semantic mappings for the beauty ontology. This is the domain's price of entry, already paid.

Cleaning and extending an existing signal

Re-parsing, de-duplicating and extending a field that was produced once and never curated.

EvidenceDeclared openly: the buy/hold/sell field covers only 33.4% of the crypto rows (1,198,479 of 3,588,282 — the other 66% read `Not mentioned`) and carries parsing leaks (`buy (implied)`, `hold/sell`). Tolid can clean and extend it. We would rather sell you the fix than hide the defect.

Scoring

Scored dimensions

Explainable, evidence-based dimensions (0–100). The radar shows the investment axes.

-

See dimension details
SpecificityRarityVolumeTraining ValueBuyer DemandEvidence StrengthData Orientation

Evidence

Dataset evidence & lineage

What the typed evidence proves the company holds — reframed for clarity and set against the market.

-

Marketplace

Dataset details

Geographic coverage

Global (multilingual sources)

Time range

2023-10 → 2025-09 (crypto) · patents and beauty: see each dataset

Delivery

Secure extract download, or scoped access

Formats

parquet, csv, jsonl

License

Derived works (annotations, graphs, metadata) are transferable. Source press text is not.

Personal data

No PII

· licence· final price on request

There is no public price grid for these corpora, and we will not manufacture one. Our own valuation engine returns wide, UNANCHORED ranges — dispersion up to ×7 on the crypto asset across three identical runs — and a dedicated deep-research study (July 2026, 103 agents, 86 claims → 25 adversarially verified) confirmed WHY: for an LLM-annotated news archive sold as annotations-only, no defensible market comparable exists. The published editor↔LLM deals disclose amounts but never volumes, so no per-document price can be derived; and the closest academic comparable (Financial PhraseBank) anchors label PEDIGREE, never a number. The gap is in the market, not in our research. So we price ON REQUEST, against two things we can defend: a measured production cost (from real cloud billing) and a label quality validated against external ground truth. The biomedical-patents asset — the strongest of the three — is under its own valuation study; its price will be shown only once anchored on real dataset transactions, never on a SaaS-subscription proxy. You do not buy the bundle: each dataset is priced on its own.

Detailed schema & sample available on access request.

Want this data?

Request access — we broker a secure deal room. Operator-reviewed, no automatic sharing.

Share this opportunity

This listing was generated automatically from public signals. It is not verified, and we are not affiliated with this company.

Coverage

Scanned sources

https://tolid.iodiscovered

Deliverable

Premium dataset report

A teaser is generated for each dataset opportunity. The full report is available on unlock.

Teaser is public · premium is locked behind access.

From the marketplace

Explore live data opportunities

Browse datasets by sector & use-case