Speaking of test suites, LLMs cannot quite curl straight quotes correctly. My Markdown editor uses the following suite: https://repo.autonoma.ca/repo/keenquotes/tree/HEAD/src/test/... Am curious whether SOTA LLMs can curl the ambiguous cases.
A Markdown-based test suite
11–15 of 15 posts
Re: A Markdown-based test suite
#12Re: A Markdown-based test suite
#13Speaking of test suites, LLMs cannot quite curl straight quotes correctly. My Markdown editor uses the following suite: https://repo.autonoma.ca/repo/keenquotes/tree/HEAD/src/test/... Am curious whether SOTA LLMs can curl the ambiguous cases.
Is curl a synonym of grok?
There are straight quotes: ' or "
and there are curly quotes: ‘ and ’ or “ and ”
"curling" then just means "figuring out whether it's an opening or closing quote"
Re: A Markdown-based test suite
#14Anybody remember "cram"? From about 10 years ago, https://bitheap.org/cram/ basically a markdown syntax (making heavy use of code-blocks) for documenting and writing tests as "shell commands and expected output" (with a bunch of the sharp edges filed off, like line endings and partial matches.) Was particularly good for easy-to-write, easy-to-review tests of unix utilities. (It's the kind of thing that you only stumb…
I get the sense that "expect tests", also known as "inline snapshots tests", are becoming more and more popular. I think that many, including myself, first learned about it from the Jane Street blog post "What if writing tests was a joyful experience?" (https://blog.janestreet.com/the-joy-of-expect-tests/) because I keep seeing references to it. Indeed, the blog post points to Cram as prior art in this space. I also think Ian Henry's article "My Kind of REPL" (https://ianthehenry.com/posts/my-kind-of-repl/) makes a very convincing case for why expect tests are useful.
Jane Street has published their `ppx_expect` library for OCaml and you can use it today: https://github.com/janestreet/ppx_expect. For other languages, here's an incomplete list of similar libraries I'm vaguely aware of:
- Rust: expect-test, (https://github.com/rust-analyzer/expect-test, quite old), insta (https://insta.rs/) and k9 (https://github.com/boujeepossum/k9)
- JavaScript: Jest has inline snapshot testing: https://jestjs.io/docs/snapshot-testing#inline-snapshots
- Python: inline-snapshot (https://15r10nk.github.io/inline-snapshot/latest/). Also discussed in this post from Pydantic: https://pydantic.dev/articles/inline-snapshot
- Go: https://github.com/hexops/autogold. I like it and use it daily at work. It has some shortcomings around working with strings as the expected value, so I'm fiddling with my own library that will be much more similar to `ppx_expect`.
- Zig: TigerBeetle has an internal `snaptest.zig` library in their codebase (https://github.com/tigerbeetle/tigerbeetle/blob/588123f219f1...), also discussed in their blog post "Snapshot Testing For The Masses" (https://tigerbeetle.com/blog/2024-05-14-snapshot-testing-for...).
Re: A Markdown-based test suite
#15The "test anything protocol" was a text based system for writing tests, I think perl might still use it I remember using it to implement the test suite for a Shakespeare language interpreter. Fun times. https://en.wikipedia.org/wiki/Test_Anything_Protocol
The markdown based tests described is more of an analog to gerken, robotframework, etc.
there's a number of similar efforts. turning english-like specs for a test into executable tests has seens lots of variants over the years but LLMs have reaaly stepped up the possibilities