Get the App
SLTechnology News&Howtos  ›  IT Information  › 

Copyright problem is difficult to solve, OpenAI is accused of illegally using book data to train AI system

Shulou Source: shulou.com Published: 2023-11-24 17:48:15 10月02日 Update

On the morning of June 30, Beijing time, it was reported that two authors sued OpenAI in the US Federal Court in San Francisco, claiming that OpenAI abused his work and used it to train ChatGPT.

Paul Tremblay and Mona Awad, writers from Massachusetts, said ChatGPT infringed the author's copyright by copying and extracting data from a large number of books without permission.

The training of advanced AI systems requires a large number of data materials, which is faced with many legal challenges. For example, source code owners point the finger at OpenAI and Microsoft's GitHub, and visual artists sue AI tools such as Stability AI, Midjourney and DeviantArt. The defendant, on the other hand, believed that the system used the copyrighted works reasonably.

When a user gives a prompt to ChatGPT, AI responds quickly, although the response is controversial. ChatGPT has been open for only two months and the number of active users reached 100 million in January.

ChatGPT and other generative AI systems use huge amounts of data to create content, much of it from the Internet. Writers Paul Tremblay and Mona Awad believe that books are key data materials because they are examples of high-quality long writing.

The complaint estimates that OpenAI's training data contains at least 300000 books, many of which are copyrighted books obtained illegally without permission.

The two plaintiffs claimed that ChatGPT could make a very accurate summary of their books, meaning that their books were incorporated into the database.

Tags: Data books systems training works copyright writers authors materials users United States precision two that is books Internet advanced key model Apple Docker Huawei Linux macOS MariaDB Microsoft MySQL NVidia OPPO Reno Redmi Linux Shulou Tech Info NVidia Xiaomi