ChatGPT’s mathematical abilities are significantly below those of an average mathematics graduate student 


health transformation institute


Joaquim Cardoso MSc
Founder, Chief Researcher & Editor
February 1, 2023


SOURCE: arxiv


Mathematical Capabilities of ChatGPT


arxiv

Simon Frieder, Luca Pinchetti, Ryan-Rhys Griffiths, Tommaso Salvatori, Thomas Lukasiewicz, Philipp Christian Petersen, Alexis Chevalier, Julius Berner


31 Jan 2023


The authors investigated the mathematical capabilities of ChatGPT by testing it on publicly available datasets, as well as hand-crafted ones, and measuring its performance against other models trained on a mathematical corpus, such as Minerva. 


The authors also tested whether ChatGPT can be a useful assistant to professional mathematicians by emulating various use cases that come up in the daily professional activities of mathematicians (question answering, theorem searching). 


  • In contrast to formal mathematics, where large databases of formal proofs are available (e.g., the Lean Mathematical Library), current datasets of natural-language mathematics, used to benchmark language models, only cover elementary mathematics. 

  • The authors address this issue by introducing a new dataset: GHOSTS. 

  • It is the first natural-language dataset made and curated by working researchers in mathematics that
    (1) aims to cover graduate-level mathematics and 
    (2) provides a holistic overview of the mathematical capabilities of language models. 

  • The authors benchmark ChatGPT on GHOSTS and evaluate performance against fine-grained criteria. 
  • The authors make this new dataset publicly available to assist a community-driven comparison of ChatGPT with (future) large language models in terms of advanced mathematical comprehension. 

The authors conclude that contrary to many positive reports in the media (a potential case of selection bias), ChatGPT’s mathematical abilities are significantly below those of an average mathematics graduate student. 


The results show that ChatGPT often understands the question but fails to provide correct solutions. 

Hence, if your goal is to use it to pass a university exam, you would be better off copying from your average peer!





ORIGINAL PUBLICATION






https://arxiv.org/pdf/2301.13867.pdf

Total
1
Shares
Deixe um comentário

O seu endereço de e-mail não será publicado. Campos obrigatórios são marcados com *

Related Posts

Subscribe

PortugueseSpanishEnglish
Total
1
Share