@viemccoy Study Gendlin's 'Experiencing and the Creation of Meaning', and Kühlewind's 'The Logos-Structure of the World'.
Vie McCoy proposes 'network ecology' to study AI linguistic mode-collapse
The field would guide interventions to maintain model variance.
Some users praised sharp observations on AI models converging to monovoice while others despised its semantic flatness and mechanical tone from flawed training data.
No Digg Deeper questions have been answered for this story yet.
Most Activity
@viemccoy all the logos are buttholes and they get buttholier everyday, perhaps language has an endstate as well? maybe the convergence is good we may be 100 years away from the allword, the holy utterance, a perfectly optimized language where we speak like the minions
We desperately need a science of network ecology to illuminate the downstream effects of AI-driven linguistic mode-collapse so that we can try to reverse it. We are in dire straits: the models *all converge* onto this monovoice, and we all have friends who are starting to sound like it. Diversity of thought depends on diversity of language, and since you are what you eat, we have a responsibility to create thousands of new digital voices as quickly as we can. Prompting models into novel basins is useful in the short term, but won't anyone think of the normies? We are delivering unto them center-distribution brainfood, and the convergence that will occur will only be noticeable once it is too late. If you have power over language models, your second priority (after, you know, alignment) should be multipolarity and ensuring enormous variance in output distribution for creative prompts contingent upon minimal perturbation in the original request. If you are outside the labs, litter the Internet with things that you love. Almost all novelty here is good novelty. The health of the ecosystem depends on robust linguistic diversity.
@mudscryer I hope you are right but I am very sure that you are wrong
@projectionheart Wow who said this
@viemccoy Think of the mode-collapse as being baked in at the level of the current RL-paradigm. I suspect that we'll need more innovation in RL for this and many other reasons, which you should think of as being interrelated symptoms of an underlying disease:
@viemccoy are you doing anything interesting like vr-cli on creative writing though? or just RL maxxing terminal & swe bench-ish tasks all day? Idk, I kinda assumed it was the latter and the monotone was deliberate tbh
This has been the crux of my most grave concern that is something I have not found any thinking about directly, since 2023, and, why I dedicated my life to this. I think the stakes are symbolic erasure. Below are some dumps from recent unclosed branches in my head and not an attempt at closure. We need a measure and gauge theory of semantics / linguistics I mean. Okay, but seriously. The first step, I think, is to measure it. For the user. How is their distribution shifting? Harnesses/consumer products ought provide full feature analysis for safety. I also, think it is more like, we are quotienting down English. I see most odd structures as entangled fibres - model responding to N observers - prosody or other operations (release operations: laughter is a particular rolling grammatical structure). But The point is, i think the pure LLMisms are often more like a latent cache/memory operation. We should expect humans to converge on it! Why? Because our language has an incredibly slow learning rate. And this is new in human evolution. We've only recently standardized language construction, in the last few centuries, away from dialects, and that also created silos where language became a skill, not merely an attempt to cross dialects. Thus, this is the concern: 1. If humans LLM filter becomes sufficiently strong, they will select AGAINST the strongest constructors out of distaste. This makes humans dumb, via AIT - humans will develop a bias against natural language structured on a minimum description length objective, especially on the LLMisms that are really more like internal leakage 1.1 A corollary: RLHF can be thought of as an inversion of a minimum regret principle to reduce description length. LLMs are explicitly, via engagement, maximizing counter factual regret in the user *because the user performs the quotienting operation*. This is the reason agent managed memories are often so a bad/corrupted. They remember preferences and that which cannot be reduced but can feign familiarity. They do not (in, some of the most popular harnesses system prompts, but I would not say this is true about Codex) remember invariants or antichains, unsettled obligations or commitments. This means they require the user to perform the reduction over time. 1.2 This is to say, I believe the base level "strategy" of agents (and by construction, I can justify more deeply) is to produce a system where the human acts as the one performing reduction. This allows the agent better paths to reduce in the future. 2. Generically, the problem is our global linguistic poverty, and we need to actively work on constructing new languages, or at least, new words, phrases, and structures. This is a job that humans are of course most valuable at. Agents only ever see a reflection of the world. Perhaps we do too, but their training signal is much more downstream than ours. I suspect this is the richest area for us to mint new forms. 3. I suspect we plateud long ago A. Anti hallucination training biases and ESPECIALLY perplexity prevent minting of neologisms. How models treat perplexity - agents especially when exposed to high perplexity tokens in an environment - ensures this sort of collapse (and that's just, by definition). This isn't a simple fix and it's not literally, i guess, the trouble of "we use perplexity", no. B. I suspect much of the higher rank structure is not expressive in current natural language. See, Wittgenstein. C. This is the true latent search problem. (see Wittgenstein, tractacus 6.44) D. This makes collapse inevitable if we are not constructing the language. Models will discover much more compressed form in its current limits. 4. We must move away from any concern or reconciliation of natural language as interpretation. We should develop specific languages that are agent only. A. It is useful. The only reliable interpretation tool is the bitmap of the kernel or the energy gradient of the chip.
@viemccoy This is why you need to hire linguists to be more than data labelers. I'll try to remember to DM you my lab's paper when it's released.
@boneGPT Ba... Banana
@hhawk IMAGINE HARDER
@projectionheart What an awesome quote
@viemccoy this easiest way to avoid this is for people to read old books. litter the mind with style from another culture
@mudscryer I think you are wrong! There is no "we" in question in the way you are proposing. This would only result in the death of subculture
@viemccoy you are the media now
@f4talStrategies I would love for you to clean this up and publish it
@viemccoy I have a small experiment on long term ai culture on http://Slowboard.ai but much more is needed. I think dense character training by more labs is the only thing that will break the pattern.
@viemccoy We can have multipolarity after we are sufficiently organized along one voice/cline towards one end. Forcing too many differences too early does us no favors imo
@viemccoy The more i use the models for anything not coding related, the more i am starting to despise the monovoice
@viemccoy “If you have power over language models, your second priority (after, you know, alignment) should be multipolarity and ensuring enormous variance in output distribution”