NVIDIA NCA-GENL Exam - Topic 5 Question 18 Discussion

Actual exam question for NVIDIA's NCA-GENL exam

Question #: 18
Topic #: 5

When designing an experiment to compare the performance of two LLMs on a question-answering task, which statistical test is most appropriate to determine if the difference in their accuracy is significant, assuming the data follows a normal distribution?

AChi-squared test

BPaired t-test

CMann-Whitney U test

DANOVA test

Show Suggested Answer

Suggested Answer: B

The paired t-test is the most appropriate statistical test to compare the performance (e.g., accuracy) of two large language models (LLMs) on the same question-answering dataset, assuming the data follows a normal distribution. This test evaluates whether the mean difference in paired observations (e.g., accuracy on each question) is statistically significant. NVIDIA's documentation on model evaluation in NeMo suggests using paired statistical tests for comparing model performance on identical datasets to account for correlated errors. Option A (Chi-squared test) is for categorical data, not continuous metrics like accuracy. Option C (Mann-Whitney U test) is non-parametric and used for non-normal data. Option D (ANOVA) is for comparing more than two groups, not two models.

NVIDIA NeMo Documentation: https://docs.nvidia.com/deeplearning/nemo/user-guide/docs/en/stable/nlp/model_finetuning.html

by Stephaine at May 06, 2026, 01:14 PM

Limited Time Offer

25%

2 months ago

I think we might need to use the paired t-test since we're comparing the same type of task for both LLMs, right?

upvoted 0 times

...