The biggest drawback of LLM is that it never answers with "I don't know" (unless it is some quote) and it just brings bullshit hallucinations which human has to reject as wrong. Thus it is mostly useless for anything serious. Personally I use it to beautify some text, but still have to do a bit of correction to fix b/s or missed context.
Post train one LLM to please another LLM, that rates the quality of the first model's responses and calls it on any bullshit! (And vice versa.)