As much as I'm against all GenAI as I view it as theft of the work of normal artists and people, I can forgive thumbnails that enhance a true image of what the video is about and just add some light effects and the likes. The one I have a massive problem with is the thumbnail of mud hut one....it's just a lie. That archway is just a made up image that doesn't appear in the final build at all.
I set all videos on youtube that use fake thumbnails to ignore channel and I've always enjoyed Matt's work so I hope this is just a settings mistake that allows google to try and force their AI rubbish onto video makers without consent.
As much as I'm against all GenAI as I view it as theft of the work of normal artists and people, I can forgive thumbnails that enhance a true image of what the video is about and just add some light effects and the likes. The one I have a massive problem with is the thumbnail of mud hut one....it's just a lie. That archway is just a made up image that doesn't appear in the final build at all.
I set all videos on youtube that use fake thumbnails to ignore channel and I've always enjoyed Matt's work so I hope this is just a settings mistake that allows google to try and force their AI rubbish onto video makers without consent.
I'm curious why would AI generated thumbnails be any different than any other AI generated image? It still took that same stolen work to create an AI generated thumbnail.
No necessarily.
Generative AI is/was a thief - of original art - yes, but not so much when it comes to product images which are often loosely copyrighted by their original companies so they can be displayed easily without problems.
Modern generative art isn't anywhere close to what it used to be either. It's still done a "diffusion" model (starts black and works backwards, so think like a sculptor in 2D).
The old way was to glue bits of other pictures together - modern ones like Dal-E 2.0 are much more original and that, in a way is scarier.
I use AI for my project - but ONLY when I can't find an artist in my budget who DOESN'T - and those people are getting scarce. I know how much art is worth (for the work I do) and the prices are either too high (because the artists aren't sufficiently skilled in real terms) OR too low because they use AI and won't admit to it.
As a writer I felt the sting but the generative art bubble is running out of rope as the tech-bros are finding just how expensive it is to run the GPUs necessary to run it.
As of this morning I just finished V0,2 of my advanced RAG system (written largely by an AI with me pulling it apart) but it can already scan simple English documents (at about 35 tokens/S on my hardware) and answer questions based on their content; intelligently! It's not quite ready for release but at some point I'll have it scan Matt's site so we can easily get answers, faster.
Take everything I say with a pinch of salt, I might be wrong and it's a very *expensive* way to learn!
No necessarily.
Generative AI is/was a thief - of original art - yes, but not so much when it comes to product images which are often loosely copyrighted by their original companies so they can be displayed easily without problems.
Modern generative art isn't anywhere close to what it used to be either. It's still done a "diffusion" model (starts black and works backwards, so think like a sculptor in 2D).
The old way was to glue bits of other pictures together - modern ones like Dal-E 2.0 are much more original and that, in a way is scarier.
I use AI for my project - but ONLY when I can't find an artist in my budget who DOESN'T - and those people are getting scarce. I know how much art is worth (for the work I do) and the prices are either too high (because the artists aren't sufficiently skilled in real terms) OR too low because they use AI and won't admit to it.
As a writer I felt the sting but the generative art bubble is running out of rope as the tech-bros are finding just how expensive it is to run the GPUs necessary to run it.
As of this morning I just finished V0,2 of my advanced RAG system (written largely by an AI with me pulling it apart) but it can already scan simple English documents (at about 35 tokens/S on my hardware) and answer questions based on their content; intelligently! It's not quite ready for release but at some point I'll have it scan Matt's site so we can easily get answers, faster.
If you use diffusion to get the result, then that result has to be based on work that was already done. Seeing as Google uses the same source as every other frontier model, aka the internet for scraping, there is no reason to believe that they have some kind of training set that doesn't use copyrighted material that Google does not have the right to use.
At least, until proven otherwise, it has to be assumed at this point that they use unlicensed content just like all the others.
Generative AI is a thief. That never stopped.
Also, please don't datamine the site and feed into your theft machine.
Err. No it hasn’t.
The images my system generates are 100% original. There’s difference between copying the style and characters of an author (although E L James did nicely out of Twilight) and that’s where it’s a problem.
Paywalls exist to stop Google and others spidering content.
I have, however, yet to learn of someone with congenital blindness painting a van Gough or even a Jackson Pollock.
I have a very specific style that I designed for my magazine that isn’t drawn by any artists anywhere.
We use the same visual language- in my case it’s called Ink and Wash but the characters are my designs drawn in that style.
and this is a public site. Nothing here is copyright except for Matt’s work. Or are you going to tell me that I’m wrong about that too?
While the corpus is much larger, there’s little to separate artists.
I have a MUCH bigger argument with the likes of Pixabay and others for facilitating a huge loss in revenue for small, often community run sites to offer stuff for artists to swap.
My “theft machine” like most of my work will be freely available for anyone under a FOSS licence so people can use and modify it.
Because knowledge that I received freely is worth paying forward.
Copyright used to be a wonderful idea. And then the corporations got involved.
Take everything I say with a pinch of salt, I might be wrong and it's a very *expensive* way to learn!
Err. No it hasn’t.
The images my system generates are 100% original. There’s difference between copying the style and characters of an author (although E L James did nicely out of Twilight) and that’s where it’s a problem.
Paywalls exist to stop Google and others spidering content.
I have, however, yet to learn of someone with congenital blindness painting a van Gough or even a Jackson Pollock.
I have a very specific style that I designed for my magazine that isn’t drawn by any artists anywhere.
We use the same visual language- in my case it’s called Ink and Wash but the characters are my designs drawn in that style.
and this is a public site. Nothing here is copyright except for Matt’s work. Or are you going to tell me that I’m wrong about that too?
While the corpus is much larger, there’s little to separate artists.
I have a MUCH bigger argument with the likes of Pixabay and others for facilitating a huge loss in revenue for small, often community run sites to offer stuff for artists to swap.
My “theft machine” like most of my work will be freely available for anyone under a FOSS licence so people can use and modify it.
Because knowledge that I received freely is worth paying forward.
Copyright used to be a wonderful idea. And then the corporations got involved.
Did you train the models from scratch on only your original work? I doubt it as no individual can produce enough training material to train any model to get something usable and even if you somehow did produce enough data, I'd like to see what datacenter you use to train said models.
These technologies are not within a normal reach of any one person, but is accessible to those with vast resources. Using any openly trained model runs the risk of it being trained on material that was not licensed to be used for training. That's just a fact. Because you, a private individual, simply do not have the resources to train these models yourself in a timely manner.
And to make something very clear; AI models don't make original work. They find the average of any vector space and output it. If anything, you are receiving average work back that you are more than likely able to make yourself much better.
Paywalls exists sure, yet several companies have been found pirating content they couldn't get their hands on for free in other ways so, what good do those paywalls do?
Your "congenital blindness" claim makes no sense at all.
And just because a site is public does not mean that the content of that site is freely available to use for training, commercial usage or otherwise. As an example; Anything you posted on Stackoverflow has a license attached to the produced work.
Seeing as no license exists to explicitly permit you to datamine the content of a site, the default copyright falls on DIY Perks as the owner of the forum. Yet another thing that web crawlers, and apparently you, don't respect.
I think we're talking past each other because you're responding to claims I never made.
You appear to have assumed that because I mentioned AI, I must have trained a foundation model on scraped data. That isn't what RAGamuffin is.
RAGamuffin is a local Retrieval-Augmented Generation (RAG) engine. It doesn't train foundation models, modify model weights or create new models. It indexes documents, retrieves relevant passages and supplies those passages as context to an LLM chosen by the user. Those are different technologies serving different purposes.
You're also conflating three separate issues:
-
how a foundation model was trained;
-
whether using such a model is lawful; and
-
whether indexing documents for search is lawful.
Those are separate technical and legal questions.
As for my comment about Matt's site, the point was made quite clearly in the final line of my original post:
"It's not quite ready for release but at some point I'll have it scan Matt's site so we can easily get answers, faster."
That's describing a search and retrieval tool. It is not a statement that I intend to train a model on the site's content or ignore permissions. Those are assumptions you've made, not claims I've made.
One of RAGamuffin's design goals is attribution. It returns the relevant source passages together with citations so the original material can be verified. The intention is to make it easier to find and credit original sources, not to obscure them.
For the avoidance of doubt, I've also created and cryptographically signed a Git snapshot of the version being discussed here. If anyone is interested in the implementation rather than assumptions about it, there will be a verifiable code snapshot to examine.
By and large, my own software and hardware is released as Free and Open Source Software, pro bono publico.
I don't work for OpenAI or any other AI company, and I'm not here to defend the AI industry in general. I was simply describing a piece of software I'm writing. If people want to discuss its actual architecture, that's fine. If the discussion is about broader objections to generative AI, that's a different conversation, and not one I'm going to pursue here.
Take everything I say with a pinch of salt, I might be wrong and it's a very *expensive* way to learn!
@omniowl I only mean I can forgive thumbnails that are using AI to effectively add a built in photoshop effect on a photo photo the creator has taken - like adding lightning between some hardware images for example. It's a pointless use of AI as you can just use photoshop but permissible under what my idea of artist theft is. Less permissible under it being a pointless waste of resources but that's a seperate AI discussion.
I think we're talking past each other because you're responding to claims I never made.
You appear to have assumed that because I mentioned AI, I must have trained a foundation model on scraped data. That isn't what RAGamuffin is.
RAGamuffin is a local Retrieval-Augmented Generation (RAG) engine. It doesn't train foundation models, modify model weights or create new models. It indexes documents, retrieves relevant passages and supplies those passages as context to an LLM chosen by the user. Those are different technologies serving different purposes.
You're also conflating three separate issues:
how a foundation model was trained;
whether using such a model is lawful; and
whether indexing documents for search is lawful.
Those are separate technical and legal questions.
As for my comment about Matt's site, the point was made quite clearly in the final line of my original post:
"It's not quite ready for release but at some point I'll have it scan Matt's site so we can easily get answers, faster."
That's describing a search and retrieval tool. It is not a statement that I intend to train a model on the site's content or ignore permissions. Those are assumptions you've made, not claims I've made.
One of RAGamuffin's design goals is attribution. It returns the relevant source passages together with citations so the original material can be verified. The intention is to make it easier to find and credit original sources, not to obscure them.
For the avoidance of doubt, I've also created and cryptographically signed a Git snapshot of the version being discussed here. If anyone is interested in the implementation rather than assumptions about it, there will be a verifiable code snapshot to examine.
By and large, my own software and hardware is released as Free and Open Source Software, pro bono publico.
I don't work for OpenAI or any other AI company, and I'm not here to defend the AI industry in general. I was simply describing a piece of software I'm writing. If people want to discuss its actual architecture, that's fine. If the discussion is about broader objections to generative AI, that's a different conversation, and not one I'm going to pursue here.
You said, and I quote:
The images my system generates are 100% original.
generates is used here, which is unambigious in this context. That has nothing to do with your RAG. You also said:
Modern generative art isn't anywhere close to what it used to be either. It's still done a "diffusion" model (starts black and works backwards, so think like a sculptor in 2D).
Doesn't matter. It's using stolen assets in the vast majority of publicly available tools and models via diffusion. And in this particular context, YouTube would be making those AI thumbnails which definitely is using assets they don't have rights to. Full stop.
Again, nothing to do with your RAG. Then again, you say this:
I use AI for my project - but ONLY when I can't find an artist in my budget who DOESN'T - and those people are getting scarce.
That does not excuse you. It is an explanation, it is not an excuse. Since you feel the need to defend the usage here my guess is that you don't always use your miraculous "100% originally generated art" machine because if you did, there'd be no reason to defend the usage, would there?
You then finish by saying:
As of this morning I just finished V0,2 of my advanced RAG system (written largely by an AI with me pulling it apart)
You also let AI write your code which I'm sure you'll try to tell me "is different than art" like so many others before you.
So when you say:
I think we're talking past each other because you're responding to claims I never made.
I don't think we are. I think you are carrying yourself poorly in this conversation and is now starting to construct what you think the conversation should be about, not what it was about.
Guys. Come on now. This is a forum to discuss the million methods of building things, not to point fingers at each other for using AI. @OmniOwl, even if you don't agree with it, @marcdraco using generative AI to code something or create magazine images is their business. Heck, I think you'd be hard-pressed to find someone who hasn't used AI in their coding now--it's kind of a no-brainer when it comes to time saving.
It also sounds to me as though all this RAGamuffin machine is doing is being a convenient search engine for this forum--and what's wrong with that? This place already has nearly 4000 posts; with a search engine powered by a more advanced search technology, there won't be any need to repost solutions to engineering problems, because the originals will be much easier to find. Marcdraco never said they were training it on all of our data: it's just compiling.
and this is a public site. Nothing here is copyright except for Matt’s work. Or are you going to tell me that I’m wrong about that too?
I think but don't quote me on this that designs and other original content are automatically under some kind of copyright protection as soon as you post them anywhere. I don't think it stretches to casual forum chatting though.
Lastly, I personally don't agree with making AI generated art or video content or whatever, but the thing is so omnipotent now that getting angry with individual people isn't really the answer. Yes, the models were trained unlawfully. But there's no point now accusing people of it, unless you're going to accuse millions of other people of doing it.
My worktable is a site of frequent detonations. Please knock so I don't start a Lithium fire when you startle me.
Guys. Come on now. This is a forum to discuss the million methods of building things, not to point fingers at each other for using AI. @OmniOwl, even if you don't agree with it, @marcdraco using generative AI to code something or create magazine images is their business.
As one of the admins here, it is very much my business when it happens on this site, to get involved with this. There is a no tolerance policy about this technology on the DIY Perks Discord server I manage.
On top of that, the context of the post is; Has Matt turned on AI Thumbnail Generation on his videos or is it a case of YouTube doing it automatically, as other creators have also experienced and thus unintentional?
If it's the former a lot of people on the Discord server will be cross as that goes against the very essence of what DIY Perks is; Beautifully human made creations.
You don't seem to have read and understood that context before making your comment.
Yes, the models were trained unlawfully. But there's no point now accusing people of it, unless you're going to accuse millions of other people of doing it.
That "but" invalidates your entire stance. I am accusing millions of other people of using this terrible technology. No question.


