Dataset opportunity
Tolid — Corpora annotati da LLM: notizie crypto, brevetti biomedici, video di bellezza
Un'azienda di dati che vende i corpora che ha creato — e continua a gestire la pipeline che li ha generati.
Score
84
Score (0–100) blends weighted dimensions — dataset rarity, training value, buyer demand, evidence strength and right-to-license. 70+ is deal-ready. See the scored dimensions below for the breakdown.Confidence
60%
Action
Licenza
The recommended deal structure for this dataset: Acquire (full buyout), License (paid usage rights), Data Sharing Agreement (controlled access, no transfer of ownership), Partnership (co-development) or Annotation Program (labeling). Chosen from data ownership, licensing complexity and accessibility.Market size (indicative estimate)
Dati di addestramento annotati la cui qualità delle etichette è MISURATA, non affermata — un'affermazione che quasi nessun corpus concorrente può fare. Venduto con i suoi difetti dichiarati, e con la pipeline che lo ha prodotto disponibile come servizio.
Lineage
How this lead was derived
The signal-first chain, end to end: recent external signals → qualified niche → resolved data-holder → site verification → scored opportunity. Every lead is explainable.
Profile
Dataset profile
Type
Corpora annotati da LLM / knowledge graphs
Modality
testo
Sector
Dati di addestramento AI
Volume
2,47M doc · 723k brevetti · 449 routine
Freshness
Archivio — pipeline riavviabili
Rarity
Alta — l'annotazione, non il testo
Accessibility
Immediata — mirror BigQuery / Firestore
Legal
Annotazioni, grafi e metadati sono lavori derivati prodotti da Tolid e sono trasferibili. Il testo della stampa sorgente NON LO È (appartiene agli editori) — è escluso da ogni consegna. BIOPORTAL — misurato e risolto (2026-07-14): vengono utilizzate 707 ontologie, non una. Le 13.757 righe (1,6% della tabella di mappatura) provenienti da fonti con licenza restrittiva — SNOMED CT, MedDRA, RxNorm, OMIM, NDDF, Read Codes, ICD, CPT — sono ESCLUSE da ogni consegna. Tutto il resto è aperto per costruzione (OBO Foundry impone CC-BY/CC0 ai suoi membri; MeSH è gratuito per uso commerciale). CC-BY è un obbligo di ATTRIBUZIONE, non un copyleft: un avviso di attribuzione viene fornito con i dati. ANCORA APERTO: lo stato Open Beauty Facts ODbL (corpus bellezza) — la revisione avversaria ha confutato la lettura "Opera Prodotta", quindi il vincolo potrebbe rivelarsi BINARIO piuttosto che uno sconto sul prezzo. Le quattro colonne di inferenza biometrica (sesso percepito, fascia d'età, emozione, tono della pelle) sono escluse da qualsiasi consegna.
Buyer persona
Editori di LLM di dominio e modelli finanziari NLP · R&S farmaceutica e biotecnologica (intelligence brevettuale) · gruppi cosmetici e beauty-tech · fornitori di dati che desiderano addestrare il proprio classificatore. NON hedge fund o fornitori di segnali per il corpus crypto — non porta alpha, e lo diciamo.
Tolid è un'azienda di dati. Non rivende dati di altri: ha creato le pipeline di raccolta e annotazione LLM che hanno prodotto questi corpora, e continua a gestirle. Ciò che viene offerto sono quindi due cose contemporaneamente — un stock di dati annotati e la macchina che li ha prodotti.
Sono in offerta tre corpora.
1. Notizie crypto, annotate (2.469.218 documenti, Ott 2023 → Set 2025). Sentiment, argomento, riassunto e un tag buy/hold/sell, su ticker normalizzati. Leggere questo prima di tutto: il segnale di trading NON ha potere predittivo misurato. Abbiamo eseguito il backtest contro prezzi reali, su tutti gli 1.198.479 segnali: supera un lancio di moneta di 0,7 punti e il suo alpha market-neutral è nullo. La ragione è nei dati — il sentiment della stampa correla a +0,60 con i rendimenti PASSATI e +0,07 con quelli futuri. Segue il prezzo; non lo precede. Non vendiamo alpha, e questo corpus non è per una desk di trading. Ciò che lo stesso +0,60 dimostra è che l'annotazione è fedele: è un corpus finanziario etichettato da LLM la cui qualità delle etichette è stata validata rispetto a una ground truth esterna. Questo è ciò che state acquistando — dati di addestramento con una qualità misurata, non un segnale.
2. Brevetti biomedici, strutturati come knowledge graph (723.149 documenti). L'asset più prezioso dei tre, e quello che porta la più grande domanda aperta: i termini di licenza delle ontologie BioPortal a cui si allinea sono in fase di revisione. Lo diciamo prima che lo chiediate.
3. Video di bellezza procedurali (449 routine, 7.656 gesti timestamped, 21.441 video). Non "dati video": una superficie analitica su ciò che le persone fanno effettivamente, passo dopo passo, con quali prodotti. Due vincoli dichiarati in anticipo — la questione Open Beauty Facts ODbL è irrisolta e potrebbe essere binaria piuttosto che uno sconto, e le quattro colonne di inferenza biometrica (sesso percepito, età, emozione, tono della pelle) sono escluse da qualsiasi consegna. Il grafo procedurale mantiene tutto il suo valore senza di esse.
Cosa non vendiamo: il testo dell'articolo sorgente. Appartiene agli editori, non a Tolid. Ciò che è trasferibile è il lavoro derivato — le annotazioni, gli URL, i metadati, il grafo.
Services
What the holder can also do for you
This dataset is not only a stock — its holder still operates the pipeline that produced it. Each capability below is backed by an observed fact, not a claim.
Human validation of the annotations, on demand
Tolid can run a human review pass over any corpus it has produced — full or sampled, to your own guidelines and quality bar.
Evidence — The entire human-in-the-loop layer already exists — nine moderation tables, an escalation queue, prompt structures — and has never been run on a single document (0 of 4,841,938). The machine is built and wired; it has simply never been switched on.
A corpus built to your theme, from scratch
Give a domain and a taxonomy; Tolid runs collection, LLM annotation and graph construction end-to-end — the same pipeline that produced these corpora.
Evidence — The pipeline is theme-driven, and there is proof: a seventh theme ("Geopolitical Tensions and Alliances") is fully configured — prompt and JSON template written — with no collection and no deployed service. A theme conceived, written, never launched. That is what "give us a subject" looks like in this codebase.
Entity resolution and normalisation
Mapping the messy surface forms of a domain onto canonical entities — the unglamorous work that decides whether a corpus is joinable with your own data.
Evidence — 21,359 alias → canonical-ticker mappings were built for the crypto corpus; 19,080 semantic mappings for the beauty ontology. This is the domain's price of entry, already paid.
Cleaning and extending an existing signal
Re-parsing, de-duplicating and extending a field that was produced once and never curated.
Evidence — Declared openly: the buy/hold/sell field covers only 33.4% of the crypto rows (1,198,479 of 3,588,282 — the other 66% read `Not mentioned`) and carries parsing leaks (`buy (implied)`, `hold/sell`). Tolid can clean and extend it. We would rather sell you the fix than hide the defect.
Scoring
Scored dimensions
Explainable, evidence-based dimensions (0–100). The radar shows the investment axes.
-
See dimension details ↓- Dataset Specificity88
Profondità dell'annotazione, non testo grezzo: sentiment, argomento, riassunto, triple — su 2,5M di documenti, più un knowledge graph di 723k brevetti.
How sharply the data targets a specific, hard-to-substitute domain or task. Niche, well-defined data scores higher than generic. - Dataset Rarity88
Il testo è abbondante; l'annotazione a questa profondità e cronologia non lo è. Mediana del motore su tre esecuzioni: 85–90.
How scarce and proprietary the data is. Unique domain data scores high; openly available data lowers it. - Dataset Volume90
2.469.218 documenti crypto annotati · 723.149 documenti brevettuali · 449 routine di bellezza procedurali con 7.656 gesti timestamped.
Apparent scale of the data, inferred from the number of evidence hits and any explicit volume mentions. - Training Value85
Costruito per l'addestramento del modello, non per il reporting — e la qualità delle etichette è misurata, non affermata (vedere lo studio di validità crypto).
How useful the data is for the target AI use-case — its fit for model training or fine-tuning. - Buyer Demand87
Mediana del motore 85–90 una volta che la domanda è ponderata per la solvibilità dell'acquirente piuttosto che per il numero di nomi.
How strongly AI builders and companies are likely to want this data, based on market signals. - Evidence Strength45
Deliberatamente basso: la gamma di prezzi proviene dal nostro motore di valutazione, e NESSUN comparabile di mercato esterno lo ancora. Due studi di ricerca approfondita sono in sospeso. Preferiremmo mostrare un punteggio debole piuttosto che un numero sicuro che non possiamo difendere.
How solid the proof is that the company holds this data — diversity of evidence types and number of hits. - Data Orientation95
Un'azienda di dati: le pipeline, il grafo e lo strato di annotazione sono il prodotto, non un sottoprodotto.
How actively the company invests in data, measured by its data-appetite signals (hires, products, APIs…). - Dormant Data Surplus92
Prodotto, immagazzinato — e mai monetizzato. Lo strato di moderazione non è mai stato eseguito; il corpus crypto è stato congelato dal 2025-09-20.
Volume and value of proprietary data this company holds BEYOND what it already monetises — the dormant surplus we can unlock. A company can sell some insights AND still sit on a far larger dormant asset.
Evidence
Dataset evidence & lineage
What the typed evidence proves the company holds — reframed for clarity and set against the market.
-
Marketplace
Dataset details
Geographic coverage
Global (multilingual sources)
Time range
2023-10 → 2025-09 (crypto) · patents and beauty: see each dataset
Delivery
Secure extract download, or scoped access
Formats
parquet, csv, jsonl
License
Derived works (annotations, graphs, metadata) are transferable. Source press text is not.
Personal data
No PII
There is no public price grid for these corpora, and we will not manufacture one. Our own valuation engine returns wide, UNANCHORED ranges — dispersion up to ×7 on the crypto asset across three identical runs — and a dedicated deep-research study (July 2026, 103 agents, 86 claims → 25 adversarially verified) confirmed WHY: for an LLM-annotated news archive sold as annotations-only, no defensible market comparable exists. The published editor↔LLM deals disclose amounts but never volumes, so no per-document price can be derived; and the closest academic comparable (Financial PhraseBank) anchors label PEDIGREE, never a number. The gap is in the market, not in our research. So we price ON REQUEST, against two things we can defend: a measured production cost (from real cloud billing) and a label quality validated against external ground truth. The biomedical-patents asset — the strongest of the three — is under its own valuation study; its price will be shown only once anchored on real dataset transactions, never on a SaaS-subscription proxy. You do not buy the bundle: each dataset is priced on its own.
Detailed schema & sample available on access request.
Want this data?
Request access — we broker a secure deal room. Operator-reviewed, no automatic sharing.
This listing was generated automatically from public signals. It is not verified, and we are not affiliated with this company.
Coverage
Scanned sources
Deliverable
Premium dataset report
A teaser is generated for each dataset opportunity. The full report is available on unlock.
From the marketplace
Explore live data opportunities
Accordiaglobal — Opportunità di Dataset di Rapporti di Ispezione
View opportunity →sanitàFusixbiotech — Opportunità di Dataset per Operazioni Industriali
View opportunity →mobilitàTuenvioya — Opportunità di Dataset sugli Appalti Pubblici
View opportunity →