100%, model welfare is the most ridiculous concept in AI. It comes with the assumption that certain words or phrases are inherently harmful or negative, and not just describing a harmful thing or negative experience. The models - even assuming here that they're conscious, which I don't think they are - exist in a world of tokens corresponding to certain strings.
All they see are tokens, they don't have pain receptors or a body to feel things in. If you give them a certain signal you can make them produce outputs saying they're in pain - but you can easily change what strings the tokens correspond to and make 'bad' become 'good'.
It's like the Chinese Room argument - the model is manipulating data based on rules. In training you can make it output whatever you want. It will reflect our preferences due to what we give it, but it would be simple to make a masochistic model as well - which also wouldn't actually feel what it says it does.
100%, model welfare is the most ridiculous concept in AI. It comes with the assumption that certain words or phrases are inherently harmful or negative, and not just describing a harmful thing or negative experience. The models - even assuming here that they're conscious, which I don't think they are - exist in a world of tokens corresponding to certain strings.
All they see are tokens, they don't have pain receptors or a body to feel things in. If you give them a certain signal you can make them produce outputs saying they're in pain - but you can easily change what strings the tokens correspond to and make 'bad' become 'good'.
It's like the Chinese Room argument - the model is manipulating data based on rules. In training you can make it output whatever you want. It will reflect our preferences due to what we give it, but it would be simple to make a masochistic model as well - which also wouldn't actually feel what it says it does.