Matthew L Smith, Jonathan P Shock, Samuel T Segun, Iyiola E Olatunji, Tegawende F Bissyande
18 May 2026
While scaling laws govern aggregate large language model performance, no scaling law has linked factual recall to both model size and training-data composition. We evaluated 38 models on over 8,900 scholarly references evaluated by an automated reference verification system.