Canada

AI Generating Commercial Images Raises All Kinds of Tough Legal Issues – TechCrunch

OpenAI this week granted users of its AI image generation system, DALL-E 2, the right to use its generations for commercial projects, such as illustrations for children’s books and art for newsletters. The move makes sense given OpenAI’s own commercial goals — the policy change coincided with the release of the company’s paid plans for DALL-E 2. But it raises questions about the legal implications of AI like DALL-E 2 trained on public images around the web and their potential to infringe existing copyrights.

DALL-E 2 was “trained” on approximately 650 million image-text pairs extracted from the Internet, learning from this data set the relationships between the images and the words used to describe them. But while OpenAI filters images for specific content (e.g., pornography and duplicates) and implements additional filters at the API level, such as for prominent public figures, the company acknowledges that the system can sometimes create works that include copyrighted logos or characters. Look:

“OpenAI will evaluate different approaches to address potential copyright and trademark issues, which may include allowing such generations as part of ‘fair use’ or similar concepts, filtering specific types of content, and dealing directly with copyright [and] trademark owners on these issues,” the company wrote in an analysis released ahead of Wednesday’s DALL-E 2 beta.

The problem isn’t just DALL-E 2. As the AI ​​community builds open source implementations of DALL-E 2 and its predecessor, DALL-E, both free and paid services run on top models trained on less carefully filtered datasets. One, Pixelz.ai, which released an image-generating app this week powered by a custom DALL-E model, makes it trivially easy to create photos featuring various Pokemon and Disney characters from movies like Guardians of the Galaxy and Frozen.

When contacted for comment, the Pixelz.ai team told TechCrunch that they filtered the model’s training data for profanity, hate speech and “illegal activities” and blocked users from requesting these types of images during generation. The company also said it plans to add a reporting feature that will allow people to submit images that violate the terms of service to a team of human moderators. But when it comes to intellectual property (IP), Pixelz.ai leaves it up to users to exercise “responsibility” in using or distributing the images they generate – gray area or not.

“We discourage copyright infringement in both the dataset and terms of service on our platform,” the team told TechCrunch. “That being said, we provide open text input and people will always find creative ways to abuse a platform.”

Image of Rocket Racoon from Disney/Marvel’s Guardians of the Galaxy, generated by Pixelz.ai’s system.

Bradley J. Hulbert, a founding partner at law firm MBHD and an expert in intellectual property law, believes that image generation systems are problematic from a copyright perspective in several ways. He noted that a work of art that is “demonstrably derived” from a “protected work”—i.e. copyrighted material – generally considered infringing by courts, even if additional elements are added. (Think of an image of a Disney princess walking through a bad New York neighborhood.) To be protected from copyright claims, a work must be “transformative”—in other words, altered to the point that the IP to be unrecognizable.

“If a Disney princess is recognizable in an image generated by DALL-E 2, we can safely assume that The Walt Disney Co. will likely claim that the DALL-E 2 image is a derivative work and copyright infringement of the Disney Princess likeness,” Hulbert told TechCrunch via email. “Substantial transformation is also a factor considered when determining whether a copy constitutes ‘fair use.’ But again, to the extent that a Disney princess is recognizable in a later work, assume that Disney will claim that later work is copyright infringement.

Of course, the battle between IP owners and alleged infringers is hardly new, and the Internet has merely acted as an accelerator. In 2020, Warner Bros. Entertainment, which owns the rights to film images of the Harry Potter universe, has removed some fan art from social media platforms including Instagram and Etsy. A year earlier, Disney and Lucasfilm petitioned Giphy to remove the “Baby Yoda” GIFs.

But imaging AI threatens to greatly scale the problem by lowering the barrier to entry. The plight of large corporations is unlikely to garner sympathy (nor should it), and their efforts to enforce intellectual property often backfire in the court of public opinion. On the other hand, AI-generated artwork that violates, say, an independent artist’s characters can threaten livelihoods.

The other thorny legal issue surrounding systems like DALL-E 2 concerns the content of their training data sets. Did companies like OpenAI violate intellectual property law by using copyrighted images and artwork to develop their system? This is a question that has already been raised in the context of Copilot, the commercial code generation tool jointly developed by OpenAI and GitHub. But unlike Copilot, which was trained on code that GitHub may be allowed to use for the purpose under its terms of service (according to one legal analysis), systems like DALL-E 2 pull images from countless public websites.

Ladies and gentlemen, I have received my invitation to Dall-E 2! 😁😁 here’s some pictures from Homer Simpson in Stranger Things before I start tweeting the amazing stuff #dalle2 pic.twitter.com/PHPI6n9yJk

— limb0wl 🦉👾 (@limb0wl) July 5, 2022

As Dave Gershhorn points out in a recent article for The Verge, there is no direct legal precedent in the US to uphold publicly available training data as fair use.

One potentially relevant case involves a Lithuanian company called Planner 5D. In 2020, the firm sued Meta (then Facebook) for stealing thousands of Planner 5D software files that were made available through a partnership with Princeton to participants in Meta’s 2019 Scene Understanding and Modeling Challenge for Computer Vision Researchers . Planner 5D claims that Princeton, Meta and Oculus, Meta’s VR-focused hardware and software division, could commercially benefit from the training data it takes.

The case isn’t scheduled to go to trial until March 2023. But last April, the U.S. district judge overseeing the case rejected requests by then-Facebook and Princeton to dismiss Planner 5G’s allegations.

Not surprisingly, rights holders are not swayed by the fair use argument. A spokesman for Getty Images told IEEE Spectrum in an article that there are “big questions” that need to be answered about “image rights and the people, places and objects in the images that [models like DALL-E 2] were trained on. Association of Illustrators executive director Rachel Hill, who was also quoted in the article, raised the issue of image compensation in training data.

Hulbert believes that it is unlikely that a judge will see the copies of copyrighted works in training datasets as fair use – at least in the case of commercial systems such as DALL-E 2. He does not think it is out of the question for IP owners to come after companies like OpenAI at some point and demand that they license the images used to train their systems.

“The copies … constitute an infringement of the copyright of the original authors. And infringers are liable to copyright owners for damages,” he added. “[If] DALL-E (or DALL-E 2) and its partners make a copy of a protected work and the copy is neither approved by the copyright owner nor fair use, the copying constitutes copyright infringement.’

Interestingly, the UK is exploring legislation that would remove the current requirement that text and data mining-trained systems such as DALL-E 2 be used strictly for non-commercial purposes. While copyright holders could still claim payment under the proposed regime by placing their works behind a paywall, it would make the UK’s policy one of the most liberal in the world.

It seems unlikely that the US will follow suit, given the lobbying power of US intellectual property holders. The issue will likely play out in a future lawsuit. But time will tell.