文章摘要
AI公司批量购买稀有书籍,用高速扫描仪切掉书脊并销毁原版,通过ISBNdb匿名订购多达百万册。2022年前书籍因无AI生成内容而更受青睐。联邦法官裁定此举为合理使用,因每次仅存一份副本。Anthropic聘请前Google Books负责人获取“全世界所有书籍”。
文章总结
好的,这是根据您的要求,对原文主要内容进行的中文重述,已保留关键细节并删减了与主题无关的评论和互动内容。
标题:Hedgie (@HedgieMarkets) 的帖子
核心内容:
人工智能公司正在批量购买稀有书籍,使用高速扫描机切掉书脊进行扫描,然后将原书销毁。一项名为ISBNdb的服务可以促成高达一百万本书的订单,并保持买家匿名。2022年之前的书籍因不含AI生成文本而备受青睐。一位联邦法官裁定这种做法属于“合理使用”,理由是销毁原书意味着同一时间只存在一份副本。Anthropic公司已聘请前谷歌图书合作负责人,以获取“世界上所有的书”。
作者评论:
一位书商告诉404媒体,那些几乎已无存世副本的稀有书籍正被送入这条流水线。那些历经战争、火灾和数百年流传的书籍,正在被粉碎,只为让AI学会写一封更好的营销邮件。ISBNdb的网站甚至直言:“‘AI公司销毁两百万本书’并不是一个能博得同情的标题”,但他们仍然围绕如何悄无声息地完成这件事建立了一整套业务。他们提供保密协议作为一项服务,并指导客户将其称为“数字保存”。
作者认为,这比AI公司抓取网络、盗版图书馆或窃取音乐的行为更恶劣,因为它是不可逆的。你可以重新上传一个网站,可以重印一本畅销书,但你无法替换被粉碎用于训练数据的、仅存三本的18世纪植物学文本。而法官判定这是合法的,因此这种行为将会加速。在2026年,“我们销毁稀有书籍并提供保密协议以防人知晓”竟成为一种合法的商业模式。
补充信息(来自作者与网友的互动):
- 扫描后保存原书需要花费一点成本,而销毁则无需成本,后者因此胜出。
- 一旦物理原版消失,就没有任何东西可以用来核对数字副本,这使得销毁原始资料与抓取网站截然不同。
- 法庭文件将Anthropic的相关项目命名为“巴拿马计划”,耗资数千万,其规划文件明确表示目标是“破坏性地扫描世界上每一本书”。
- 扫描会丢失物理书籍所承载的批注、印刷工艺和装订等信息。
评论总结
根据评论内容,主要围绕AI公司购买旧书进行数字化扫描(可能涉及销毁)的行为展开讨论,观点分歧明显。以下是总结:
支持/理解方观点: - 数字化有助于保存内容,物理载体本身价值有限(评论4、17、21)。关键引用:"I think people are, on the whole, too precious about old things... it is not the paper that imbues the book with historical value"(评论4);"Pretty much anything that is still under copyright should be valuable only for its content, not for the physical medium"(评论17)。 - 这是版权法问题的结果,而非AI公司恶意(评论3、20)。关键引用:"it's more a problem with copyright law than AI companies"(评论3);"The judge made exactly the right call and the companies are following the law"(评论20)。 - 被销毁的并非真正稀有古籍,多为仍有版权的普通旧书(评论5、10、17)。关键引用:"they're not after ancient texts - they're after the books that there's still copyright on"(评论5);"probably aren't 'rare books' in the way people are thinking"(评论10)。
批评/担忧方观点: - 可能销毁稀有书籍,造成不可逆损失(评论1、3、29)。关键引用:"You can't replace the last three copies of an 18th-century botanical text once someone shreds them for training data"(评论3);"if they are destroyed in the process of becoming training data, they'll be even harder to obtain"(评论29)。 - 缺乏证据证明销毁行为,但令人不安(评论6、22、26)。关键引用:"I don't see any proof of shredding here"(评论6);"Is there any proof of this at all beyond this random message?"(评论22)。 - 应公开数字化成果,而非仅用于训练数据(评论18、30)。关键引用:"why not upload the scanned books to the internet archive while already at it?"(评论18);"People would have at least somewhat less of a problem with this if they also put up an archive of PDFs"(评论30)。
中立/质疑方观点: - 需区分真正稀有古籍与普通旧书(评论16、27)。关键引用:"Is there evidence of this? ... they could very well be describing what is only occurring to in-print or non-rare books"(评论16);"Why would you need to shred a book from the 1800s when it is in the public domain?"(评论27)。 - 这是版权法扭曲的结果,应推动法律改革(评论20)。关键引用:"The fix here is to change the law to permit training AI without destroying the original materials"(评论20)。