Is one run enough? Reproducibility of flagship large language models across temperature and reasoning settings in biomedical text processing
{{output}}
Background: To quantify run-to-run reproducibility of Gemini 3 Flash Preview and GPT-5.2 for trial-success classification across temperature and reasoning/thinking settings and determine whether single-run reporting suffices. ... ...