Dataset opportunity

Tolid — LLM 标注语料库:加密货币新闻、生物医学专利、美妆视频

一家销售其构建语料库并仍运营构建这些语料库的管道的公司。

LLM 标注语料库 / 知识图谱文本领域 LLM 训练、知识图谱、情绪模型🌍 Francetolid.io2026年7月16日

Confidence

60%

Market size (indicative estimate)

标注训练数据,其标签质量是可衡量的,而非断言的——几乎没有竞争性语料库能做出此声明。已声明其缺陷并提供生成它的管道作为服务进行销售。

Lineage

How this lead was derived

The signal-first chain, end to end: recent external signals → qualified niche → resolved data-holder → site verification → scored opportunity. Every lead is explainable.

Profile

Dataset profile

Type

LLM 标注语料库 / 知识图谱

Modality

文本

Sector

AI 训练数据

Volume

2.47M 文档 · 723k 专利 · 449 个流程

Freshness

存档 — 管道可重启

Rarity

高 — 标注,而非文本

Accessibility

即时 — BigQuery / Firestore 镜像

Legal

标注、图谱和元数据是 Tolid 生成的衍生作品,可转移。源新闻文本不可转移(属于出版商)——不包含在任何交付内容中。BIOPORTAL — 已测量和解决(2026-07-14):使用了 707 个本体集,而非一个。来自限制性许可来源的 13,757 行(占映射表的 1.6%)—— SNOMED CT、MedDRA、RxNorm、OMIM、NDDF、Read Codes、ICD、CPT — 不包含在任何交付内容中。其他所有内容均按设计开放(OBO Foundry 要求其成员遵守 CC-BY/CC0;MeSH 可免费用于商业用途)。CC-BY 是署名义务,而非 copyleft:数据附带署名通知。仍未解决:Open Beauty Facts ODbL 状态(美妆语料库)——对抗性审查驳斥了“生成作品”的解读,因此该限制可能为二元而非价格折扣。四个生物识别推断列(感知性别、年龄范围、情绪、肤色)不包含在任何交付内容中。

Buyer persona

领域 LLM 和金融 NLP 模型出版商 · 制药和生物技术研发(专利情报) · 化妆品集团和美妆科技 · 希望训练自己分类器的供应商。不适用于对冲基金或加密货币语料库的信号供应商——它不产生阿尔法,我们已明确说明。

Tolid 是一家数据公司。它不转售他人的数据:它构建了生成这些语料库的集合和 LLM 标注管道,并且至今仍在运营它们。因此,本次提供的是两样东西——一批标注好的数据,以及生成它的机器

三份语料库可供选择。

1. 加密货币新闻,已标注(2,469,218 份文档,2023 年 10 月 → 2025 年 9 月)。 包含对标准化代币的情绪、主题、摘要和买入/持有/卖出标签。请务必先阅读此内容:交易信号不具备任何可测量的预测能力。 我们自己对所有 1,198,479 个信号进行了回测,与真实价格进行比较:其表现比抛硬币好 0.7 个点,并且其市场中性阿尔法为零。原因在于数据——新闻情绪与过去回报的相关性为 +0.60,与未来回报的相关性为 +0.07。它跟随价格;它并不领先。我们不销售阿尔法,此语料库不适用于交易部门。 同样是 +0.60 的相关性证明了标注的准确性:这是一个 LLM 标记的金融语料库,其标签质量已根据外部真实情况进行了验证。您购买的是这个——具有可衡量质量的训练数据,而不是信号。

2. 生物医学专利,结构化为知识图谱(723,149 份文档)。 三者中最有价值的资产,也是承载最大未决问题的资产:与其对齐的 BioPortal 本体集的许可条款正在审查中。我们会在您询问之前告知您。

3. 程序化美妆视频(449 个流程,7,656 个带时间戳的手势,21,441 个视频)。 不是“视频数据”:而是对人们实际操作的分析表面,一步一步地使用哪些产品。提前声明两个限制——Open Beauty Facts ODbL 问题尚未解决,可能为二元而非折扣,并且四个生物识别推断列(感知性别、年龄、情绪、肤色)不包含在任何交付内容中。程序化图谱在没有这些的情况下仍保留其全部价值。

我们不销售的内容:源文章文本。它属于出版商,不属于 Tolid。可转移的是衍生作品——标注、URL、元数据、图谱。

Services

What the holder can also do for you

This dataset is not only a stock — its holder still operates the pipeline that produced it. Each capability below is backed by an observed fact, not a claim.

Human validation of the annotations, on demand

Tolid can run a human review pass over any corpus it has produced — full or sampled, to your own guidelines and quality bar.

EvidenceThe entire human-in-the-loop layer already exists — nine moderation tables, an escalation queue, prompt structures — and has never been run on a single document (0 of 4,841,938). The machine is built and wired; it has simply never been switched on.

A corpus built to your theme, from scratch

Give a domain and a taxonomy; Tolid runs collection, LLM annotation and graph construction end-to-end — the same pipeline that produced these corpora.

EvidenceThe pipeline is theme-driven, and there is proof: a seventh theme ("Geopolitical Tensions and Alliances") is fully configured — prompt and JSON template written — with no collection and no deployed service. A theme conceived, written, never launched. That is what "give us a subject" looks like in this codebase.

Entity resolution and normalisation

Mapping the messy surface forms of a domain onto canonical entities — the unglamorous work that decides whether a corpus is joinable with your own data.

Evidence21,359 alias → canonical-ticker mappings were built for the crypto corpus; 19,080 semantic mappings for the beauty ontology. This is the domain's price of entry, already paid.

Cleaning and extending an existing signal

Re-parsing, de-duplicating and extending a field that was produced once and never curated.

EvidenceDeclared openly: the buy/hold/sell field covers only 33.4% of the crypto rows (1,198,479 of 3,588,282 — the other 66% read `Not mentioned`) and carries parsing leaks (`buy (implied)`, `hold/sell`). Tolid can clean and extend it. We would rather sell you the fix than hide the defect.

Scoring

Scored dimensions

Explainable, evidence-based dimensions (0–100). The radar shows the investment axes.

-

See dimension details
SpecificityRarityVolumeTraining ValueBuyer DemandEvidence StrengthData Orientation

Evidence

Dataset evidence & lineage

What the typed evidence proves the company holds — reframed for clarity and set against the market.

-

Marketplace

Dataset details

Geographic coverage

Global (multilingual sources)

Time range

2023-10 → 2025-09 (crypto) · patents and beauty: see each dataset

Delivery

Secure extract download, or scoped access

Formats

parquet, csv, jsonl

License

Derived works (annotations, graphs, metadata) are transferable. Source press text is not.

Personal data

No PII

· licence· final price on request

There is no public price grid for these corpora, and we will not manufacture one. Our own valuation engine returns wide, UNANCHORED ranges — dispersion up to ×7 on the crypto asset across three identical runs — and a dedicated deep-research study (July 2026, 103 agents, 86 claims → 25 adversarially verified) confirmed WHY: for an LLM-annotated news archive sold as annotations-only, no defensible market comparable exists. The published editor↔LLM deals disclose amounts but never volumes, so no per-document price can be derived; and the closest academic comparable (Financial PhraseBank) anchors label PEDIGREE, never a number. The gap is in the market, not in our research. So we price ON REQUEST, against two things we can defend: a measured production cost (from real cloud billing) and a label quality validated against external ground truth. The biomedical-patents asset — the strongest of the three — is under its own valuation study; its price will be shown only once anchored on real dataset transactions, never on a SaaS-subscription proxy. You do not buy the bundle: each dataset is priced on its own.

Detailed schema & sample available on access request.

Want this data?

Request access — we broker a secure deal room. Operator-reviewed, no automatic sharing.

Share this opportunity

This listing was generated automatically from public signals. It is not verified, and we are not affiliated with this company.

Coverage

Scanned sources

https://tolid.iodiscovered

Deliverable

Premium dataset report

A teaser is generated for each dataset opportunity. The full report is available on unlock.

Teaser is public · premium is locked behind access.

From the marketplace

Explore live data opportunities

Browse datasets by sector & use-case