Assessing AI Chatbots Performance in Comprehensive Standardized Test Preparation; A Case Study with GRE

AI-generated keywords: Artificial intelligence chatbots standardized tests GRE graduate school

AI-generated Key Points

  • Research paper evaluates performance of three AI chatbots: Bing, ChatGPT, and GPT-4
  • Focuses on their abilities in addressing standardized test questions using GRE as a case study
  • Administered 137 quantitative reasoning questions and 157 verbal questions
  • Questions categorized into varying levels of difficulty
  • Evaluated chatbots' proficiency in addressing image-based questions and measured uncertainty level
  • Results show varying degrees of success across chatbots, influenced by model sophistication and training data
  • GPT-4 emerged as most proficient, particularly in complex language understanding tasks
  • Demonstrates evolution of AI in language comprehension and ability to pass exams with high scores
  • Explores chatbots' abilities in assessing verbal reasoning, quantitative reasoning, critical thinking, and analytical writing skills
  • Provides valuable insights into utilization of AI in standardized test preparation.
Also access our AI generated: Comprehensive summary, Lay summary, Blog-like article; or ask questions about this paper to our AI assistant.

Authors: Mohammad Abu-Haifa, Bara'a Etawi, Huthaifa Alkhatatbeh, Ayman Ababneh

20 Pages, 6 figures, and 6 tables
License: CC BY 4.0

Abstract: This research paper presents a comprehensive evaluation of the performance of three artificial 10 intelligence chatbots: Bing, ChatGPT, and GPT-4, in addressing standardized test questions. Graduate record examination, known as GRE, serves as a case study in this paper, encompassing both quantitative reasoning and verbal skills. A total of 137 quantitative reasoning questions, featuring diverse styles and 157 verbal questions categorized into varying levels of difficulty (easy, medium, and hard) were administered to assess the chatbots' capabilities. This paper provides a detailed examination of the results and their implications for the utilization of artificial intelligence in standardized test preparation by presenting the performance of each chatbot across various skills and styles tested in the exam. Additionally, this paper explores the proficiency of artificial intelligence in addressing image-based questions and illustrates the uncertainty level of each chatbot. The results reveal varying degrees of success across the chatbots, demonstrating the influence of model sophistication and training data. GPT-4 emerged as the most proficient, especially in complex language understanding tasks, highlighting the evolution of artificial intelligence in language comprehension and its ability to pass the exam with a high score.

Submitted to arXiv on 26 Nov. 2023

Ask questions about this paper to our AI assistant

You can also chat with multiple papers at once here.

AI assistant instructions?

Results of the summarizing process for the arXiv paper: 2312.03719v1

This research paper presents a comprehensive evaluation of the performance of three artificial intelligence chatbots: Bing, ChatGPT, and GPT-4. The study focuses on their abilities in addressing standardized test questions and uses the Graduate Record Examination (GRE) as a case study. This exam encompasses both quantitative reasoning and verbal skills and is commonly used by universities to evaluate applicants to their graduate programs. A total of 137 quantitative reasoning questions and 157 verbal questions were administered to assess the chatbots' capabilities. These questions were categorized into varying levels of difficulty. The study also evaluated the chatbots' proficiency in addressing image-based questions and measured their uncertainty level. The results revealed varying degrees of success across the chatbots, highlighting the influence of model sophistication and training data. GPT-4 emerged as the most proficient, particularly in complex language understanding tasks. This demonstrates the evolution of artificial intelligence in language comprehension and its ability to pass exams with high scores. In addition to evaluating the chatbots' performance on standardized test questions, this paper explores their abilities in assessing verbal reasoning, quantitative reasoning, critical thinking, and analytical writing skills – all crucial factors for graduate school applicants. Overall, this research provides valuable insights into the utilization of artificial intelligence in standardized test preparation.
Created on 05 Jan. 2024

Assess the quality of the AI-generated content by voting

Score: 0

Why do we need votes?

Votes are used to determine whether we need to re-run our summarizing tools. If the count reaches -10, our tools can be restarted.

The previous summary was created more than a year ago and can be re-run (if necessary) by clicking on the Run button below.

Similar papers summarized with our AI tools

Navigate through even more similar papers through a

tree representation

Look for similar papers (in beta version)

By clicking on the button above, our algorithm will scan all papers in our database to find the closest based on the contents of the full papers and not just on metadata. Please note that it only works for papers that we have generated summaries for and you can rerun it from time to time to get a more accurate result while our database grows.

Disclaimer: The AI-based summarization tool and virtual assistant provided on this website may not always provide accurate and complete summaries or responses. We encourage you to carefully review and evaluate the generated content to ensure its quality and relevance to your needs.