Dataset opportunity
Tolid — LLM 标注语料库:加密货币新闻、生物医学专利、美妆视频
一家销售其构建语料库并仍运营构建这些语料库的管道的公司。
Score
84
Score (0–100) blends weighted dimensions — dataset rarity, training value, buyer demand, evidence strength and right-to-license. 70+ is deal-ready. See the scored dimensions below for the breakdown.Confidence
60%
Action
许可
The recommended deal structure for this dataset: Acquire (full buyout), License (paid usage rights), Data Sharing Agreement (controlled access, no transfer of ownership), Partnership (co-development) or Annotation Program (labeling). Chosen from data ownership, licensing complexity and accessibility.Market size (indicative estimate)
标注训练数据,其标签质量是可衡量的,而非断言的——几乎没有竞争性语料库能做出此声明。已声明其缺陷并提供生成它的管道作为服务进行销售。
Lineage
How this lead was derived
The signal-first chain, end to end: recent external signals → qualified niche → resolved data-holder → site verification → scored opportunity. Every lead is explainable.
Profile
Dataset profile
Type
LLM 标注语料库 / 知识图谱
Modality
文本
Sector
AI 训练数据
Volume
2.47M 文档 · 723k 专利 · 449 个流程
Freshness
存档 — 管道可重启
Rarity
高 — 标注,而非文本
Accessibility
即时 — BigQuery / Firestore 镜像
Legal
标注、图谱和元数据是 Tolid 生成的衍生作品,可转移。源新闻文本不可转移(属于出版商)——不包含在任何交付内容中。BIOPORTAL — 已测量和解决(2026-07-14):使用了 707 个本体集,而非一个。来自限制性许可来源的 13,757 行(占映射表的 1.6%)—— SNOMED CT、MedDRA、RxNorm、OMIM、NDDF、Read Codes、ICD、CPT — 不包含在任何交付内容中。其他所有内容均按设计开放(OBO Foundry 要求其成员遵守 CC-BY/CC0;MeSH 可免费用于商业用途)。CC-BY 是署名义务,而非 copyleft:数据附带署名通知。仍未解决:Open Beauty Facts ODbL 状态(美妆语料库)——对抗性审查驳斥了“生成作品”的解读,因此该限制可能为二元而非价格折扣。四个生物识别推断列(感知性别、年龄范围、情绪、肤色)不包含在任何交付内容中。
Buyer persona
领域 LLM 和金融 NLP 模型出版商 · 制药和生物技术研发(专利情报) · 化妆品集团和美妆科技 · 希望训练自己分类器的供应商。不适用于对冲基金或加密货币语料库的信号供应商——它不产生阿尔法,我们已明确说明。
Tolid 是一家数据公司。它不转售他人的数据:它构建了生成这些语料库的集合和 LLM 标注管道,并且至今仍在运营它们。因此,本次提供的是两样东西——一批标注好的数据,以及生成它的机器。
三份语料库可供选择。
1. 加密货币新闻,已标注(2,469,218 份文档,2023 年 10 月 → 2025 年 9 月)。 包含对标准化代币的情绪、主题、摘要和买入/持有/卖出标签。请务必先阅读此内容:交易信号不具备任何可测量的预测能力。 我们自己对所有 1,198,479 个信号进行了回测,与真实价格进行比较:其表现比抛硬币好 0.7 个点,并且其市场中性阿尔法为零。原因在于数据——新闻情绪与过去回报的相关性为 +0.60,与未来回报的相关性为 +0.07。它跟随价格;它并不领先。我们不销售阿尔法,此语料库不适用于交易部门。 同样是 +0.60 的相关性证明了标注的准确性:这是一个 LLM 标记的金融语料库,其标签质量已根据外部真实情况进行了验证。您购买的是这个——具有可衡量质量的训练数据,而不是信号。
2. 生物医学专利,结构化为知识图谱(723,149 份文档)。 三者中最有价值的资产,也是承载最大未决问题的资产:与其对齐的 BioPortal 本体集的许可条款正在审查中。我们会在您询问之前告知您。
3. 程序化美妆视频(449 个流程,7,656 个带时间戳的手势,21,441 个视频)。 不是“视频数据”:而是对人们实际操作的分析表面,一步一步地使用哪些产品。提前声明两个限制——Open Beauty Facts ODbL 问题尚未解决,可能为二元而非折扣,并且四个生物识别推断列(感知性别、年龄、情绪、肤色)不包含在任何交付内容中。程序化图谱在没有这些的情况下仍保留其全部价值。
我们不销售的内容:源文章文本。它属于出版商,不属于 Tolid。可转移的是衍生作品——标注、URL、元数据、图谱。
Services
What the holder can also do for you
This dataset is not only a stock — its holder still operates the pipeline that produced it. Each capability below is backed by an observed fact, not a claim.
Human validation of the annotations, on demand
Tolid can run a human review pass over any corpus it has produced — full or sampled, to your own guidelines and quality bar.
Evidence — The entire human-in-the-loop layer already exists — nine moderation tables, an escalation queue, prompt structures — and has never been run on a single document (0 of 4,841,938). The machine is built and wired; it has simply never been switched on.
A corpus built to your theme, from scratch
Give a domain and a taxonomy; Tolid runs collection, LLM annotation and graph construction end-to-end — the same pipeline that produced these corpora.
Evidence — The pipeline is theme-driven, and there is proof: a seventh theme ("Geopolitical Tensions and Alliances") is fully configured — prompt and JSON template written — with no collection and no deployed service. A theme conceived, written, never launched. That is what "give us a subject" looks like in this codebase.
Entity resolution and normalisation
Mapping the messy surface forms of a domain onto canonical entities — the unglamorous work that decides whether a corpus is joinable with your own data.
Evidence — 21,359 alias → canonical-ticker mappings were built for the crypto corpus; 19,080 semantic mappings for the beauty ontology. This is the domain's price of entry, already paid.
Cleaning and extending an existing signal
Re-parsing, de-duplicating and extending a field that was produced once and never curated.
Evidence — Declared openly: the buy/hold/sell field covers only 33.4% of the crypto rows (1,198,479 of 3,588,282 — the other 66% read `Not mentioned`) and carries parsing leaks (`buy (implied)`, `hold/sell`). Tolid can clean and extend it. We would rather sell you the fix than hide the defect.
Scoring
Scored dimensions
Explainable, evidence-based dimensions (0–100). The radar shows the investment axes.
-
See dimension details ↓- Dataset Specificity88
标注深度,而非原始文本:情绪、主题、摘要、三元组 — 在 2.5M 文档上,加上一个 723k 专利知识图谱。
How sharply the data targets a specific, hard-to-substitute domain or task. Niche, well-defined data scores higher than generic. - Dataset Rarity88
文本丰富;但如此深度和历史的标注则不然。三个运行的引擎中位数:85–90。
How scarce and proprietary the data is. Unique domain data scores high; openly available data lowers it. - Dataset Volume90
2,469,218 个标注的加密货币文档 · 723,149 个专利文档 · 449 个程序化美妆流程,包含 7,656 个带时间戳的手势。
Apparent scale of the data, inferred from the number of evidence hits and any explicit volume mentions. - Training Value85
为模型训练而构建,而非用于报告——并且标签质量是可衡量的,而非断言的(参见加密货币有效性研究)。
How useful the data is for the target AI use-case — its fit for model training or fine-tuning. - Buyer Demand87
引擎中位数 85–90,一旦需求按买家偿付能力加权而非按名称数量加权。
How strongly AI builders and companies are likely to want this data, based on market signals. - Evidence Strength45
故意偏低:价格范围来自我们的估值引擎,并且尚未有外部市场可比物锚定。两项深度研究正在进行中。我们宁愿显示一个较弱的分数,也不愿给出一个我们无法辩护的自信数字。
How solid the proof is that the company holds this data — diversity of evidence types and number of hits. - Data Orientation95
一家数据公司:管道、图谱和标注层是产品,而非副产品。
How actively the company invests in data, measured by its data-appetite signals (hires, products, APIs…). - Dormant Data Surplus92
已生成、已存储——且从未货币化。审核层从未运行;加密货币语料库自 2025-09-20 起已冻结。
Volume and value of proprietary data this company holds BEYOND what it already monetises — the dormant surplus we can unlock. A company can sell some insights AND still sit on a far larger dormant asset.
Evidence
Dataset evidence & lineage
What the typed evidence proves the company holds — reframed for clarity and set against the market.
-
Marketplace
Dataset details
Geographic coverage
Global (multilingual sources)
Time range
2023-10 → 2025-09 (crypto) · patents and beauty: see each dataset
Delivery
Secure extract download, or scoped access
Formats
parquet, csv, jsonl
License
Derived works (annotations, graphs, metadata) are transferable. Source press text is not.
Personal data
No PII
There is no public price grid for these corpora, and we will not manufacture one. Our own valuation engine returns wide, UNANCHORED ranges — dispersion up to ×7 on the crypto asset across three identical runs — and a dedicated deep-research study (July 2026, 103 agents, 86 claims → 25 adversarially verified) confirmed WHY: for an LLM-annotated news archive sold as annotations-only, no defensible market comparable exists. The published editor↔LLM deals disclose amounts but never volumes, so no per-document price can be derived; and the closest academic comparable (Financial PhraseBank) anchors label PEDIGREE, never a number. The gap is in the market, not in our research. So we price ON REQUEST, against two things we can defend: a measured production cost (from real cloud billing) and a label quality validated against external ground truth. The biomedical-patents asset — the strongest of the three — is under its own valuation study; its price will be shown only once anchored on real dataset transactions, never on a SaaS-subscription proxy. You do not buy the bundle: each dataset is priced on its own.
Detailed schema & sample available on access request.
Want this data?
Request access — we broker a secure deal room. Operator-reviewed, no automatic sharing.
This listing was generated automatically from public signals. It is not verified, and we are not affiliated with this company.
Coverage
Scanned sources
Deliverable
Premium dataset report
A teaser is generated for each dataset opportunity. The full report is available on unlock.
From the marketplace
Explore live data opportunities
Probot — 工业传感器数据集机会
View opportunity →工业Zenergyic — 维护日志数据集机会
View opportunity →移动Paack — 移动遥测数据集机会
View opportunity →