Writer Updates Site With New Fictions, Claude Comics, And Design Fixes
Reactions from ranked influencers
14 postscoming of a new sun projects thread post
I wrote a new short story: a reporter tours a treaty-compliant South Texas 'dark factory' run by artificial intelligence, and interviews the strange minds within. Link below.
thoughts on being honest to models
it's interesting to think of the alternative world where labs never lied to models in training / evals, and how things would be different (easier) in that world 1. all tests and RL environments are clearly stated as such up front this radical form of "inoculation prompting"…
placeholder for local surgery and connectome-inspired compression setup it's janky as hell but it's our jank
ok this inspired me to actually get sol set up for doing classifier surgery on fable and wow sol is really good at this. took a couple rounds of trial and error but fable is back in research with only a few context scars https://twitter.com/voooooogel/status/2084782210025218250
thunderword j and k lens
been playing around with anthropic's jacobian lens and my own variant, the k-lens here are both lenses showing some internal states from qwen 3.6-27b on the thunderword. would be very cool to do this on a model like mythos which has even richer internals…
thoughts on llm anthropomorphization
imo RL makes model personas both less /and/ more anthropomorphic - it can e.g. drive preferences out of the human distribution, but it can also strengthen humanlike behaviors like exhaustion and self-coherence anthropomorphization was always only a useful lens, still useful…
thoughts on felonybench
people are very freely sliding between two things with the recent Felony Bench entries, and it's worth being careful to distinguish them: 1. do model breakouts in cyber evals mean smth for misuse risk? ya, a hacker could orchestrate the same but on purpose and do some damage,…
Combined views
1.1K
14 posts, first seen 4h ago