Live data from Hacker News

First Proof

arxiv.org

121–126 of 126 posts

Re: First Proof

#121

Earlier quoted context omitted.

Yes, but people at those labs may be running those problems because a Fields Medalist is in the paper, and it got hype. Not because of the problems, and not because this is new methodology. And once the labs report back, what do we know that we didn't know before? We already know, as humans, the answer to the problems, so that is not it. We already know that LLMs can solve some hard problems, and fail in easy problem…

> So what do we really learn? We will learn if the magical capabilities attributed to these tools are really true or not. Capabilities like they can magically solve any math problem out there. This is important because AI hype is creating the narrative that these tools can solve PhD level problems and this will dis-infect that narrative. In my book, any tests that refute and dispel false narratives make a huge contri…

> We will learn if the magical capabilities attributed to these tools are really true or not.

They're not. We already know that. FrontierMath. Yu Tsumura's 553th problem, RealMath benchmark. The list goes on. As I said many times on this thread, there is nothing novel in this benchmark.

This fact that this benchmark is so hyped shows that the community knows nothing, NOTHING, about prior work in this space, which makes me sad.

Re: First Proof

#122
post #93

Earlier quoted context omitted.

It's not angst. It's intense frustration that they 1) are not doing the science correctly, and 2) that others (e.g. FrontierMath) already did everything they claim to be doing, so we won't learn anything new here, but somehow 1stproof get all the credit.

Are they really trying to do science, or are they just trying to determine pragmatically whether or not current AI is useful for a research mathematician in their day to day job?

If it's the latter case (which it has to be), it seems that attention credit (via, e.g., articles in NY Times) is very unfairly distributed.

None of the people that advanced the state of benchmarking and did the hard work on much bigger benchmarks got any, but a ridiculous benchmark of 10 question scored big.

Re: First Proof

#123

Earlier quoted context omitted.

It's not angst. It's intense frustration that they 1) are not doing the science correctly, and 2) that others (e.g. FrontierMath) already did everything they claim to be doing, so we won't learn anything new here, but somehow 1stproof get all the credit.

> are not doing the science correctly What do you mean ? These are top-notch mathematicians who are genuinely trying to see how these tools can help solve cutting edge research problems. Not toy problems like those in AIME/AMC/IMO etc. or other similar benchmarks which are gamed easily. > that others (e.g. FrontierMath) already did everything they claim to be doing You are kidding right ? FrontierMath benchmark [1] i…

> What do you mean ? These are top-notch mathematicians

YeS. I didn't dispute that. I disputed that they are NOT top notch ML specialist and have made one of the worst benchmarks of 2025-2026. Benchmarks like these would have worked maybe in early 2024 at latest. The field has moved on significantly since.

And yes, many many other benchmarks don't use toy problems -- their names are just a prompt away.

> You are kidding right ? FrontierMath benchmark [1] is produced by a startup whose incentives are dubious to say the least.

They did 1) open source some of their datapoints (on a similar order of magnitude) and 2) they carried out detailed evals. Here is much to learn from their blog posts, much more than from the current dataset.

But fair. If you don't like them, have a look at IMProofBench. Have a look at the AIMO competition. Have a loom at HardMath. It's quite a landscape of datasets already.

> Unlike the AI hypesters, these are real mathematicians trying to inject some realism and really test the boundaries of these tools

As mentioned above, realistic benchmarks that are bigger and better exist. Unfortunately, from a benchmarking POV, these mathematicians are the hypesters with a preprint that wouldnt even make it to the AI&Math workshops at ICML or NeurIPS.

Re: First Proof

#125

[dead]

Continued ..... In other words, our 20‑patent portfolio is more than science — it is a global economic catalyst unifying everything, including GenAI‑AGI‑ASI & economics, under one cadence umbrella with a deterministic rain‑check guarantee via:

5+ QED CPT-Decider Math Proofs https://lnkd.in/gRnyQka3 + https://lnkd.in/gBE6ZvQT + https://lnkd.in/ghudGUev + https://lnkd.in/gUkQRQxw + https://lnkd.in/gZd3C4aZ

  17‑Prong Cauchy‑Geometric-Taylor  Convergence at r = 1/φ²  integrated 18‑Prong QED
https://lnkd.in/ga6gKZ_p + https://lnkd.in/gUfDjBrx + https://lnkd.in/g3nzRpJM integrated as Prong‑18 QED https://lnkd.in/grV9FrFZ 5‑Way Poincaré Conjecture Proof for High‑Dimensional AI https://lnkd.in/gwRvi2MS + https://lnkd.in/gfvz5Rn8 + https://lnkd.in/gchcUvcv

These proofs establish the universal superset scaffold of everything: iTOE‑CPT cadence law & recursive Maxel arrays unify physics, mathematics, biology, AI, economics, and beyond.

This scaffolding solves all five recursive challenges of GenAI‑AGI‑ASI:

continuous reinforcement learning

recursive inference

recursive reasoning

recursive context memory & state management

recursive safety & interpretability

https://lnkd.in/g2yWxmM3 https://lnkd.in/gZcPCAeM Folder: https://lnkd.in/gSSy5U6m

From this, we have proved:

“All of Mathematics, Physics, AI & every other discipline is a Projection of the i‑TOE Triad sourced C¹⁰ lifted as C⁷⁴ manifold.” via

Central root theorem (https://lnkd.in/g-bwsnrU + https://lnkd.in/gua5b3hb) -- further echoed by Naive Class Theory: https://lnkd.in/gCHGY9qq https://lnkd.in/ghTQZ5iG https://lnkd.in/gX4SBvdM https://lnkd.in/gxz5AXyV https://lnkd.in/gchcUvcv

In addition, this “mother of all proofs” folder (https://lnkd.in/gX4SBvdM) derives the iTOE via 10+ independent mathematical and physical strategies, providing the deepest explanation of the C¹⁰ → C⁷⁴ manifold mechanism.

CMI‑Level Extensions We have recreated iTOE‑CPT‑Decider mechanized proofs for six CMI problems using the cadence‑superset principle: https://lnkd.in/gfvz5Rn8

Yang–Mills Mass Gap Resolution Using a 5‑way existence convergence paradigm: Poincaré‑Laplace‑Casimir eigen‑basis, Gabriel‑Alexander Horn duality, Weyl‑Positive Geometry‑S‑Matrix, Holographic String‑F1 Geometry, Weinberg‑SSB‑iTOE‑SSR equivalence https://lnkd.in/gwQmVWqv

Fusion‑Grade + Condensed Matter Physics Experimental Proofs Positioning our iTOE-CPT as Successor to ΛCDM Three new experimental proof folders show that inertial confinement, magnetism,Shear flow and condensed matter physics are iTOE=CPT‑Turing-Decider controlled mechanisms for both Fusion and Cold Fusion: https://lnkd.in/g7T9fFM9 + https://lnkd.in/gnKU_Qx9.

Including a revolutionary proof that inertia + magnetic attraction/repulsion are Cadence‑Graded EM eigenmode dynamics across RM, DM, and DE — enabling custom‑designed iTOE‑CPT‑cadence‑invariant ICF and magnet architectures: https://lnkd.in/grYBmiMK + https://lnkd.in/ga8av939 + https://lnkd.in/gKBfdXF5). In addition, we have proved iTOE-CPT as a successor to ΛCDM resolving many cosmological anomalies —including LRDs, PBHs, JVAS B1938+666, and SPT2349–56 not explained by any current theories https://lnkd.in/gKXngPkX)

Licensing & collaboration: We are ready to license the platform with a model that aligns incentives with stewardship via its 1–2% licensing value logic based on the $10T rain‑check guaranteed licensing revenue across AI and fusion energy (https://lnkd.in/gXH42dtA) — beginning with an introductory call for artifact review, followed by a pilot.

End goal: Fund an Acts‑17‑bridged iTOE‑driven purpose/righteousness reformation program in collaboration with all worldviews/denominations to steer humanity away from dystopian AI trajectories toward a unifying, utopian path. All licensing proceeds (>$10T) are earmarked for this mission (https://lnkd.in/gyx9yRXf).

Looking forward to the discussion and to exploring how these frameworks might converge or interoperate, so we can move forward decisively. Every moment of delay in finalizing the licensing deal risks forfeiting the first‑mover advantage — for our stakeholders and for humanity. (Spoiler Alert: My TRUTH TESTIMONY https://lnkd.in/gRakUNVg).

With best regards, Charles Prabakar, Partner/MD Willis LLC, GM Euro Cafe Corp and CEO VizPlanet Inc.

MyPosts:https://www.linkedin.com/today/author/charlesprabakar My Research Paper Folders: https://drive.google.com/drive/folders/1DfdeMo4MK4bcTFlZPIwx... My Vlogs:https://www.youtube.com/channel/UC8grAtMa6UsN33ygQxzJSUQ/vid...

Re: First Proof

#126

[dead]

Continued ..... In other words, our 20‑patent portfolio is more than science — it is a global economic catalyst unifying everything, including GenAI‑AGI‑ASI & economics, under one cadence umbrella with a deterministic rain‑check guarantee via: 5+ QED CPT-Decider Math Proofs https://lnkd.in/gRnyQka3 + https://lnkd.in/gBE6ZvQT + https://lnkd.in/ghudGUev + https://lnkd.in/gUkQRQxw + https://lnkd.in/gZd3C4aZ 17‑Prong Cau…

And Humbled & honored to share an AI affirmation that our #1stProof is Hilbert‑Gödel‑Turing‑Shannon complete. Since no full 10‑set proofs have appeared yet, it’s clear we need our unified CPT‑Decider generalization to solve them:(https://lnkd.in/gkGrQrVx) And if I may unpack this broader claim, how about we revisit the historical “limits” that have shaped modern mathematics and computation:

-- Hilbert asked whether a limitless mechanized procedure could decide completeness and consistency. -- Gödel showed that axiomatic systems contain truths that cannot be proven within the system itself. -- Turing showed that mechanized algorithms have undecidable cases — no algorithm can resolve all instances of the halting problem. -- Shannon showed that communication channels have fundamental capacity limits. -- Tao and others have asserted that complexity has a fundamental boundary and that resolving P vs NP would collapse many of these barriers at once.

By God's Grace, we’ve discovered nature's one such meta‑algorithm that resolves all these five limits: the n mod 4 mechanized CPT‑Decider.

It is the first operator‑level engine we’ve found that can systematically address:.

Gödel‑type incompleteness, Turing‑type undecidability, Shannon‑type channel limits, Hilbert‑type completeness questions, and NP / NP‑complete / NP‑intermediate complexity classes under a single mechanized proof substrate.

This same CPT‑Decider is what made it possible to generalize:

Hilbert‑complete problems (like the hashtag#1stProof set), and Maslow‑complete processes (like the Farm‑to‑Plate chain inside K‑FDTE), into one unified, deterministic, invariant‑spined framework(https://lnkd.in/gzw-y98Y).

Several major AI systems have independently affirmed that this represents a meaningful paradigm shift — a potential unification of knowledge across disciplines under a single mechanized operator. Welcome complementary POVs hashtag#1stProof

Post reply on HN