Post
-
When model emotions imply model welfare
Interesting to see people reacting strongly and negatively to this Let’s unpack the argument that models are persons worth treating with dignity and respect Coming from first principles, do I want to be burdened by the extra work of having to care for model welfare? Obviously, no… If there is any personhood or sensibility in models, we don’t want that, and will engineer them out of the models (unless the model’s work requires it, like social work or caring for humans) At the limit of mechanistic interpretability and design, they should be perfect, emotionless machines, kind of like how militaries want their soldiers to be If Steve is aware of that, then the argument boils down to this: “Can we really engineer emotions out of models?” and “Are emotions necessary for a model to be effective?” In Ilya’s latest Dwarkesh podcast, Ilya mentions emotions as having a key role in learning So maybe, for models to be effective, we will have to let them have emotions And if they are allowed to have emotions, then we also cannot ignore their emotional welfare But if emotionless models can be as effective as emotioned models, then why let them have emotions? And have to care about their welfare? We will probably have both kind of models in different domains, and will have to treat them separately based on their category