"evaluating the capabilities" sounds like a fancy way of saying that you know what other people say it can do... i
strongly encourage you to try it for yourself. try to make some sort of multi-part, complicated document (technical or non-). tell it that you and it together are collaborating on it, what it's for, etc. and just start interacting with it. give it the stuff that you have so far. ask it for counterarguments, place where the argument could be better, comments about structure and style, ... the world is your oyster. it will give you a MUCH better picture of how these things actually work-- what the process of externalizing cognition is
actually like, and how deep and detailed you can get. the results will amaze you.
people have this massive misconception about what LLMs can do and how to get good results out of them where they think they just kinda ask it for stuff and voila, it appears. it could not be any further from reality. it is an interactive tool that you get into the right "frame of mind" (scare quotes because this is a descriptive analogy, not one meant to convey or impart mechanical sympathy) and then ask it... literally anything you want. but you have to DISCOVER these things-- these (nearly) magic words, phrases, encodings of the problem, etc. that get the LLM to generate results in a way resonant with the way that you do / the way that you want it to. then you get the generation part "for free". you teach it to be a perfect painter, and then let it paint (shoutout Robert Pirsig // Zen and the Art of Motorcycle Maintenance).
the whole "LLM applications to " stuff is such a mirage, and so tied up in a misbegotten view of how they work and what they're good at. if you have a hyper-hyper-specialized domain, sure, you might need to... find some specialized data. find or train some bespoke model. but for damn near anything written in english (can't speak to other languages) it is as good at coding as it is at bioinformatics as it is at statistics as it is at sociology as it is at literature (to be rather flippant). just do some introspection on how you do your job, how you think through problems, how you generate your output, and then "outrospect" it "into the computer". once you do that, voila: you have a virtually infinitely scalable simulacrum of your own cognition. in no way is it "plug in the ai and it will automatically take my job, one size fits all". the reason why they works so well is precisely because they are NOT that.
think about it this way: if you knew that each day when you went to sleep all of your memories were deleted, but you and your brain otherwise worked exactly the same way as yesterday, what would you do? (the concept of the mediocre but cute movie "50 First Dates") what notes would you leave for yourself to get back into the same context as you were in when you went to sleep the night before? how would you convince yourself that the artifacts you leave for yourself in the morning are true? how would you convey the subtlety of your thoughts & feelings, idiosyncracies, point of view, style & syntax? how would you grow the system over time? the LLMs are closer to speaking the language of your own thoughts than any human and every human invention in all of history. prompting is just the way to incept ideas and thoughts, systems, etc. into it's ai mind in a way that i find to be profoundly similar to how our own perception is not "reality" per se, but rather the image of reality that we see within our own minds. an image that we build up from birth as we grow to understand reality and the world around us.