I mean opensource is nice, but that's not actually going to make healthcare safer. Whatever flavour of AI needs to be deterministic, which llama, et al are not. even if you turn the temperature right down. As others have pointed out, its the training set that actually makes a model behave, hence why models are freely given away by large companies.
Why must it be deterministic? Humans are not but we're still can have trust in humans.
Moreover, if its not deterministic, how can you assess if its safe? Sure you can run many many iterations, but how do you know when its safe enough? LLMs encourage freeform entry, which means the testing space is fucking massive.
Does writing with a different syntactic style give different outcomes? Does spelling mistakes lead to increase morbidity? thats a test plan I don't want to have to run (unless you are paying me megabucks.)
Humans are not deterministic as you point out, which is why you need to control for as many unknowns as possible when testing in a health setting.