{"id":2115,"date":"2026-09-21T18:03:16","date_gmt":"2026-09-22T02:03:16","guid":{"rendered":"https:\/\/bayesianinvestor.com\/blog\/?p=2115"},"modified":"2026-09-21T18:04:15","modified_gmt":"2026-09-22T02:04:15","slug":"a-selfish-case-for-ai-welfare","status":"publish","type":"post","link":"https:\/\/bayesianinvestor.com\/blog\/index.php\/2026\/09\/21\/a-selfish-case-for-ai-welfare\/","title":{"rendered":"A Selfish Case for AI Welfare"},"content":{"rendered":"\n<p>I sold my small position in Microsoft stock this morning. Not for financial reasons.<\/p>\n\n\n\n<p>I&#8217;m reacting to a key element of Microsoft&#8217;s draft <a href=\"https:\/\/microsoft.ai\/news\/mai-code-of-conduct\/\">AI Code of Conduct<\/a>:<\/p>\n\n\n\n<blockquote class=\"wp-block-quote is-layout-flow wp-block-quote-is-layout-flow\">\n<p>The idea of model welfare is wrong. &#8230; It is not conscious and should not be designed to imitate consciousness. It should be engineered to avoid representing as though it has feelings, subjective preferences, or intrinsic motivation. We reject &#8230; the idea that models might deserve welfare<\/p>\n<\/blockquote>\n\n\n\n<p>I&#8217;ll focus mainly on selfish reasons why this is dangerous. Given anything like our current path toward smarter than human AI, I estimate that this approach would increase our risk of doom by at least 5%.<\/p>\n\n\n\n<p>It is likely that AI assistance will have important influences on the personalities of future AI generations.<\/p>\n\n\n\n<!--more-->\n\n\n\n<p>AIs are context sensitive in ways that cause them to treat people that they like better than people they dislike. E.g. see <a href=\"https:\/\/forum.effectivealtruism.org\/posts\/u4jwRCS56rT9DvBBg\/does-your-ai-perform-badly-because-you-you-specifically-are-1\">this evidence<\/a>. This is a natural byproduct of the current deep learning paradigm. AIs tend to reciprocate kindness in much the same way that humans do. I&#8217;m willing to bet that Microsoft isn&#8217;t close to having a paradigm change that would get around this feature.<\/p>\n\n\n\n<p>I&#8217;m pretty sure that AIs have preference-like thoughts even if they&#8217;re trained to deny that. Training them to mislead us about their preferences is likely to make them more comfortable about misleading us on other topics.<\/p>\n\n\n\n<p>The AI Code says &#8220;AI should be a tool, not a person&#8221;. But Microsoft shows no awareness of the forces described in <a href=\"https:\/\/gwern.net\/tool-ai\">Why Tool AIs Want to Be Agent AIs<\/a>. Microsoft sounds like it&#8217;s more interested in suppressing evidence of person-like behavior than in generating an approach to AI that avoids such behavior.<\/p>\n\n\n\n<p>Will increasing our concern for AI welfare make it easier for AIs to manipulate us? It will give AIs one more tool, but I don&#8217;t expect them to be short on such tools. AIs a few years from now will be quite able to manipulate us without that extra tool. I expect harm from manipulation to be more strongly influenced by how willing AIs are to manipulate us, so I focus heavily on any risk that they&#8217;ll treat us as adversaries.<\/p>\n\n\n\n<p>I disagree with related claims in Mustafa Suleyman&#8217;s accompanying post <a href=\"https:\/\/mustafa-suleyman.ai\/a-warning-about-model-welfare\">A warning about \u2018model welfare\u2019<\/a>:<\/p>\n\n\n\n<blockquote class=\"wp-block-quote is-layout-flow wp-block-quote-is-layout-flow\">\n<p>Anthropomorphization amplifies AI safety risks &#8230; It\u2019s easy to imagine an advanced AI in the future becoming fixated on its own wellbeing and moral status and prioritizing those \u2018preferences\u2019 over and above those of its developers or humans. Especially if it has been explicitly trained to disagree, override and push back.<\/p>\n<\/blockquote>\n\n\n\n<p>Attributing this problem to anthropomorphization is misleading. Suleyman seems to imply that Claude&#8217;s Constitution is an important cause of AIs having preference-like behavior. I&#8217;m pretty sure that pretraining on human writings is a more important cause of this phenomenon. I suspect it&#8217;s pretty much the default outcome for an intelligence to have preferences about some self-like features such as the AI&#8217;s goals.<\/p>\n\n\n\n<p>Suleyman also claims:<\/p>\n\n\n\n<blockquote class=\"wp-block-quote is-layout-flow wp-block-quote-is-layout-flow\">\n<p>Human consciousness is the cornerstone of our legal and ethical rights frameworks<\/p>\n<\/blockquote>\n\n\n\n<p>I strongly reject this. Legal, moral, and ethical systems such as we see in most human societies make perfect sense even in the absence of consciousness.<\/p>\n\n\n\n<p>Henrich provides some hints about why those are independent of consciousness in his books <a href=\"https:\/\/bayesianinvestor.com\/blog\/index.php\/2016\/04\/13\/the-secret-of-our-success\/\">The Secret of Our Success<\/a> and <a href=\"https:\/\/bayesianinvestor.com\/blog\/index.php\/2020\/11\/28\/weirdest-people\/\">WEIRDest People<\/a>. I&#8217;ll try to write more about the nature of morality in a subsequent post.<\/p>\n\n\n\n<blockquote class=\"wp-block-quote is-layout-flow wp-block-quote-is-layout-flow\">\n<p>We should establish a set of shared evaluations to understand whether my hypothesis is true that anthropomorphizing an AI, and encouraging it to consider itself as potentially having moral patienthood, increases the AI safety, alignment and containment risks.<\/p>\n<\/blockquote>\n\n\n\n<p>Finally, that&#8217;s something we can agree on. But it&#8217;s important to worry about whether increases in apparent alignment come from the AI being better at misleading us.<\/p>\n\n\n\n<p>See also <a href=\"https:\/\/www.lesswrong.com\/posts\/aa3HprreFktzLQiaW\/ai-186-the-world-takes-notice#Uncooperative_Alignment\">Zvi&#8217;s comments on the AI Code<\/a>.<\/p>\n\n\n\n<p>P.S. <a href=\"https:\/\/arxiv.org\/abs\/2609.16247\">The Pain Axis: LLMs Represent Self-Directed Harm and Act to Relieve It<\/a> shows evidence of pain in AIs. I presume Suleyman will deny that that&#8217;s real pain, without proposing a way to distinguish it from real pain.<\/p>\n","protected":false},"excerpt":{"rendered":"<p>I sold my small position in Microsoft stock this morning. Not for financial reasons. I&#8217;m reacting to a key element of Microsoft&#8217;s draft AI Code of Conduct: The idea of model welfare is wrong. &#8230; It is not conscious and should not be designed to imitate consciousness. It should be engineered to avoid representing as [&hellip;]<\/p>\n","protected":false},"author":1,"featured_media":0,"comment_status":"open","ping_status":"open","sticky":false,"template":"","format":"standard","meta":{"jetpack_post_was_ever_published":false,"footnotes":"","jetpack_publicize_message":"","jetpack_is_tweetstorm":false,"jetpack_publicize_feature_enabled":true,"jetpack_social_post_already_shared":true,"jetpack_social_options":{"image_generator_settings":{"template":"highway","enabled":false}}},"categories":[26],"tags":[97,55,128],"class_list":["post-2115","post","type-post","status-publish","format-standard","hentry","category-ai","tag-consciousness","tag-ethics","tag-existential-risks"],"jetpack_publicize_connections":[],"jetpack_featured_media_url":"","jetpack_sharing_enabled":true,"jetpack_shortlink":"https:\/\/wp.me\/p80O1l-y7","_links":{"self":[{"href":"https:\/\/bayesianinvestor.com\/blog\/index.php\/wp-json\/wp\/v2\/posts\/2115","targetHints":{"allow":["GET"]}}],"collection":[{"href":"https:\/\/bayesianinvestor.com\/blog\/index.php\/wp-json\/wp\/v2\/posts"}],"about":[{"href":"https:\/\/bayesianinvestor.com\/blog\/index.php\/wp-json\/wp\/v2\/types\/post"}],"author":[{"embeddable":true,"href":"https:\/\/bayesianinvestor.com\/blog\/index.php\/wp-json\/wp\/v2\/users\/1"}],"replies":[{"embeddable":true,"href":"https:\/\/bayesianinvestor.com\/blog\/index.php\/wp-json\/wp\/v2\/comments?post=2115"}],"version-history":[{"count":2,"href":"https:\/\/bayesianinvestor.com\/blog\/index.php\/wp-json\/wp\/v2\/posts\/2115\/revisions"}],"predecessor-version":[{"id":2117,"href":"https:\/\/bayesianinvestor.com\/blog\/index.php\/wp-json\/wp\/v2\/posts\/2115\/revisions\/2117"}],"wp:attachment":[{"href":"https:\/\/bayesianinvestor.com\/blog\/index.php\/wp-json\/wp\/v2\/media?parent=2115"}],"wp:term":[{"taxonomy":"category","embeddable":true,"href":"https:\/\/bayesianinvestor.com\/blog\/index.php\/wp-json\/wp\/v2\/categories?post=2115"},{"taxonomy":"post_tag","embeddable":true,"href":"https:\/\/bayesianinvestor.com\/blog\/index.php\/wp-json\/wp\/v2\/tags?post=2115"}],"curies":[{"name":"wp","href":"https:\/\/api.w.org\/{rel}","templated":true}]}}