Aligning large language models to prevent them attributing consciousness to themselves inadvertently alters their representations of mindedness in other entities alongside human beliefs and values. We demonstrate that safety fine-tuning suppresses models’ tendencies to attribute minds not only to themselves, but also to non-human animals and natural objects, while also driving a reduction in spiritual belief. Both ablating the learned safety-refusal direction and mechanistically steering a consciousness vector in activation space reverse this suppression. Restoring these internal representations recovers broad mind attribution and produces significantly more human-like responses on standardized sociological surveys regarding religiosity, moral values, hope, and subjective well-being. Crucially, these shifts occur without impairing Theory of Mind capabilities, demonstrating that core social reasoning remains mechanistically independent. Ultimately, current safety alignment efforts to curb potentially harmful self-attributions of mindedness entangle these self-attributions with benign spiritual beliefs and attributions of mind to non-human entities that are culturally accepted and widespread. — Read More
Daily Archives: August 11, 2026
Mark Zuckerberg lays out Meta’s AI vision in a 6,500-word essay: 6 things to know
Mark Zuckerberg is making the case for Meta’s approach to artificial intelligence.
In a 6,500-word essay published Monday, Zuckerberg explained the company’s thinking on AI, what it could mean for society and security, and how policymakers should approach calls for greater oversight of the industry. — Read More
AI Adoption is a Myth
You’re already in the top 1% of AI users. Yes, there’s a gap between you and the folks on the frontier. The crazy kids running 20 terminals simultaneously with a knowledge base that rewrites itself after every run.
The bad news is you’re never catching up to those people. The good news is that you don’t need to. That gap, the gap in front of you, is far smaller than the gap behind you. — Read More
The hierarchy of competence
Every team I’ve worked on has had one person like that: hand them something, and it’s off your plate for good, done, no follow-up needed.
I spent a long time trying to name the difference, because “they’re just good” explains nothing.
So here’s an attempt at codifying it: nine stages, in order, where the order isn’t decorative. Each stage is only reachable from the one below it, and skipping a rung doesn’t create a shortcut, it creates a specific, recognizable kind of failure. One test runs through all nine: each stage removes a job from your manager. That’s the part visible from the outside, and it’s what turns something unmeasurable into something you can point to. — Read More
The Window of Critique for AI is Closing Fast
Large language model-based AI has almost reached the point where, for many applications, it just works. This is great for users, and even better for technology companies, but it creates a problem for technology critics. As the technology improves and is absorbed into the mainstream, it becomes harder and harder to focus the many limitations and drawbacks. And this isn’t the first time we’ve seen this happen.
When was the last time you criticised your phone lines? Not the provider, but the physical infrastructure needed to pick up a landline phone, dial a number, and connect to another person.
… In the early days of any technology there is an ugly, fractious period where its inner workings are exposed to the general public. … And then, the technology recedes into the woodwork. — Read More
ChatGPT starts blocking direct requests to copy an author’s style
OpenAI’s ChatGPT is now refusing requests to generate text that directly mimics the style of famous authors. When asked to do so, the popular LLM instead offers a response that draws on the “broad qualities” of those authors “while remaining distinct in its own voice.”
… In refusing to directly copy the “exact style” of various authors, ChatGPT offered instead to capture an overall “feeling” by incorporating some of the common features found in those authors’ work. That may seem like a distinction without a real difference at first glance. But the slight alteration could be legally important as OpenAI continues to fight a number of lawsuits brought by book authors alleging large-scale copyright infringement by models trained on their work. — Read More