Live data from Hacker News

I accidentally turned LLM memory into program analysis

pwning.systems

61–70 of 94 posts

Re: I accidentally turned LLM memory into program analysis

#61
post #26

So he's using an LLM to generate data stored in an "is_a" representation. That's so classic AI. Soon, he'll discover that he needs quantifiers. Then that "for all" is too strong sometimes, and he needs "for most". That way lies Cyc. It's not a bad idea. But it does have a history.

In general, what all the big LLM providers are doing is moving towards classical & neural (neuro-symbolic) AI - even though they dont publicly admit it because that would counter their claims for years of "scale is all you need" (which has vanished with diminishing returns, see $MS / altman's GPT-5 bet).

The various advances in LLM technology tend to rhyme with the advances in computer programming in general. For example, the stunts that involved getting LLMs to create compilers and browsers are really just extremely expensive[0] versions of genetic programming (none of it would have worked without using the test-suite as a fitness-function). The recent news of migrations from one test-framework to another (featuring Asana, I believe), was something that we could always do trivially in a language that was based on S-Expressions (Lisp, Scheme, etc).

In fact, both Cyc and the "AI" Labs have the _same basic thesis_: Intelligence is, primarily, a data entry problem. They just disagree about what kinds of heuristics should be run over that data (logic-programs, neural-nets).

Whenever I read about someone using LLMs to write code, it _very closely_ resembles how Lenat was using Eurisko/Cyc to solve problems: they let the system run continuously, and they "nudge" it in "interesting" directions, "when it gets stuck", or "runs out of steam". (Quotes indicate their phrasing, not mine)

Even Lee Spector noticed something analogous with his genetic programming system. When he tried to get it to discover optimal data structures (or maybe it was sorting algorithms, I forget), the system would quickly "run out of steam", without a solution. But when they added new verbs/opcodes to the system, that were a better fit for that domain (e.g. index-based memory loads + stores), it converged on a solution very quickly (even for GP, domain specific languages keep delivering unreasonable wins). You will note that this rhymes with the "micro-theories" of Cyc, which in turn rhyme with the SLMs of the AI labs.

In my personal experience, most of the "silver bullets" do not work (obviously), but some of them do nudge you towards being a better programmer (by refining your intuition about the problem specifically, and computers more generally).

EDIT: just remembered something. LLMs tend to produce larger and larger programs over time, and most people (IIRC) interpret this as a kind of entropy. This happens to rhyme with a similarly observed behavior in GP. Most genetic programs that do not have a fitness function that rewards smaller size, tend to grow in an unbounded way. The reason for this, is that most of the code/genes are useless, and random mutations do not lobotomize the program under evolution. I suspect that the coding LLMs tend to grow their code for similar reasons.

[0]: I suspect that, this was mostly a triumph of enormous amounts of hardware, more than the actual LLM technology. I further suspect that a traditional GP approach, on the same quantity of hardware, could have gotten there faster (if not better as well).

Re: I accidentally turned LLM memory into program analysis

#62
I encountered this with trying to have LLMs populate facts about electoral campaigns. Like when a candidate drops out, when endorsements happen, but also if a candidate is un-endorsed or drops and rejoins. It also needed to handle if any of these facts were incorrect.

I settled on a knowledge graph in Postgres and downloading/storing the source documents so it could iterate on past results without more scraping or network calls.

This blog post helped me understand security analysis in this context! A lot of the important systems around malware analysis or large scale system security (the parts people really care about) clicked for me. So thanks for writing it.

Anyways I hope we can find some pattern to converge on with this wrt "facts management" since I feel this is currently something a lot of people and LLMs are struggling with. In practice current LLMs working with episodic memory feels similar to a grandparent with dementia scrawling things down in notebooks, crossing things out, and getting very confused.

Re: I accidentally turned LLM memory into program analysis

#63
post #42

Earlier quoted context omitted.

In general, what all the big LLM providers are doing is moving towards classical & neural (neuro-symbolic) AI - even though they dont publicly admit it because that would counter their claims for years of "scale is all you need" (which has vanished with diminishing returns, see $MS / altman's GPT-5 bet).

How do you know this?

LLMs using REPL are one instance of symbols to "bounce" their prediction against domain constraints for verification. Also shout out to Gary Marcus who was right after all (and LLM companies wasting 100s of billions of dollars for years in-between on pure scaling).

Re: I accidentally turned LLM memory into program analysis

#64

Datalog seems like a way to "spell" knowledge graph (KG). The article touches on Datalog statements changing over time. One ingredient I think would be good to add to the system is to make every statement carry "providence" metadata. The providence should be sufficient to enable later confirmation that a statement is still valid or if the statement needs to be reformed without the need to remake the entire graph from…

The OP rediscovered that frontier models prefer to reason over logical scaffolding for complex tasks. They excel at technical work with many constraints as they can perfectly maintain the references, flow graph, and evidence states while they works through your conformance gates.

Many people here might disagree; but they're holding it wrong. Those folks should ask:

1. Am I using free tier tokens? 2. Am I working on trivial software? 3. Am I expecting models to adapt to my ways of thinking?

Anyone affect by any of these three mistakes will maintain an impenetrable filter of perpetual ignorance about model capabilities. Since the OP came with receipts, I'm reproducing an example graph below.

From a plan in my active research project:

  Phase 0   resolve -> vendor -> pin             [USER+agent]  -> G0 CORPUS-RESOLVED   [PASS]

  Track B   B1 Abusalah              [WRITTEN]   [agent]       -> G1B Q-BOUNDARY
            B2 Guan-Riazanov-Yuan    [WRITTEN]                  (KILL-BOUNDARY risk)
            B3 Mahmoody-Smith-Wu     [WRITTEN]  + O-MSW-VERIFIER-COST closed (t = 588, MEASURED)
                                     -> G1B PASSED: BOUNDARY-ESTABLISHED

  Track A   05 Binius                [WRITTEN]   [agent]       -> G1 AUDITS-COMPLETE [PASS] 2026-08-28
            07 FRI-Binius 2024/504   [WRITTEN] — vendored by USER; calculated Q-FLOOR point obtained
            06 HOBBIT                [WRITTEN] — O-QUEUE-M escalated to 179 Phase 2
            01 Wesolowski            [WRITTEN] — hinge splits; C_overlap = 0, C_tail = 57.50% computed
            02 Tight VDFs            [WRITTEN] — no omega; O(log T) proof threads; black-box RO base impossible
            03 LaBRADOR (light)      [WRITTEN] — calculated, not measured; translation remains open
            04 LatticeFold+          -- deferred to pre-141

  Track C   [PASS] C0a contract -> C0b exact simulator -> G0C TRANSLATION-READY
            C2 Alwen-Serbinenko [TECHNIQUE-ONLY] -> C1 SoW [PASS] ANCESTRY-FOUND
                                                -> C4 Seeded PoW [FORM-ONLY]
                                                                 -> G1C [RED] WORK-LEMMA-MISSING
                                                                          [WARN] ANCESTRY-CLASSICAL
                                                                 -> L0 [PASS] SCOPE-FROZEN
                                                                      |- E0 response [PASS] PRESENT
                                                                      `- L1-L2 local [PASS] CARRIER-QUOTIENT
                                                                           -> J0 [PASS] TECHNIQUE-FIT CONFIRMED
                                                                                -> admission decision [PASS] CANONICAL RESULT
                                                                                     -> A0 [AMBER] REFERENCE-TESTED
                                                                                          -> A1 [PASS] ROOT-LEAF FEASIBILITY LEAD
                                                                                               -> semantic equivalence + epoch [USER]
            C3, C5  -- deferred to pre-141

  Phase 2   Q-TAIL / Q-FLOOR synthesis           [agent]       -> [PASS] G2: TAIL-MECHANISM-FOUND · FLOOR-OPEN
  Phase 3   reference ledger for 141              [agent]       -> CLOSED; awaits fresh USER authorization

Re: I accidentally turned LLM memory into program analysis

#68

I reached a similar conclusion: LLMs should only really sit at the terminals of request fulfilment. 1. User request understanding: natural language -> a more rigorous representation, in my case Datalog. 2. Result interpretation: facts and derived facts -> natural language. Between those terminals, the work should be mechanical reasoning over some ontology or formal knowledge structure. That connects to another princi…

What you call "Weathering" has been a constant gripe of mine. We have LLM-driven softwares toward that almost seem to start from scratch every time a request comes in - there are mechanisms to learn or generalize, like writing out a memory, but they are not reliable or reliable in general. There is no convenient lever to be able to say "yes this is in the memory but the request seems like it needs a fresh scan of data, so ignore your memory", or the opposite "you can infer this from stuff in the memory - don't re-reason!". There is some work like Dynamic Cheatsheets [1] and Agentic Context Engineering [2] that have studied this aspect, but we are far from a generally reliable solution. And till we have that, I think the system variations for systems trying to solve this problem are going to be (a) LLM-leaning: create unstructured memory files, with human in the loop as a filter to reject inaccurate responses (b) LLM-as-a-layer: what you describe and the article kind of is doing.

[1] https://aclanthology.org/2026.eacl-long.333/ [2] https://openreview.net/pdf?id=eC4ygDs02R

Re: I accidentally turned LLM memory into program analysis

#69

Datalog seems like a way to "spell" knowledge graph (KG). The article touches on Datalog statements changing over time. One ingredient I think would be good to add to the system is to make every statement carry "providence" metadata. The providence should be sufficient to enable later confirmation that a statement is still valid or if the statement needs to be reformed without the need to remake the entire graph from…

if you aren't familiar with the literature, then good on you for getting the right insight. provenance and support are staples of the datalog world.

the other really handy thing here is that it very straightforward to differentiate datalog rule sets. so if you change a fact, you can run a smaller solution that tells you which consequences are affected by the change. under the assumption that the evaluator is pure, we don't need to keep the support graph as you suggest, we can just generate the deltas from evolution of the differentiated form.

detailed provenance can be of real application utility, but if you just care about the accounting, keeping track of the number of supports for each fact is sufficient.

another model which is fun is to make the version of the database (monotonic time) an explicit field in your base facts, assuming you can afford to keep the whole history. a deletion then is just a negative-support, and you can ask questions about the state of the world at any time.

Re: I accidentally turned LLM memory into program analysis

#70

Datalog seems like a way to "spell" knowledge graph (KG). The article touches on Datalog statements changing over time. One ingredient I think would be good to add to the system is to make every statement carry "providence" metadata. The providence should be sufficient to enable later confirmation that a statement is still valid or if the statement needs to be reformed without the need to remake the entire graph from…

I think you mean provenance, not providence
Post reply on HN