Building the Blueprint for AI-Powered Compliance Checking: Analyzing ChatGPT-4 & Gemini by Question Category in Engineering Regulations
摘要
This paper forms a part of an in progress PhD research, exploring the transformative potential of question-and-answer (Q&A) models within the context the law that regulats the practice of engineering professions (law) in the Kingdom of Bahrain. Driven by recent advancements in Machine Learning (ML) and large language models like ChatGPT-4 and Gemini. Q&A systems offer exceptional promise to improve access to crucial regulatory information and empower informed decision-making for engineers. The researchers delve into the history of Q&A models and the key role of word transformation and embeddings in their accuracy. By analyzing data pertaining to regulatory requirements outlined in the aforementioned law, the researchers conducted a comprehensive comparison of Chat GPT-4 and Gemini focusing on five user-relevant question types: factual, critical thinking, hypothetical, open-ended, and comparison. The performance analysis of ChatGPT-4 and Gemini across different question types indicates that both models performed effectively in factual questions. ChatGPT-4 achieved 93.6% correct answers, with 4.3% partially correct, while Gemini achieved 89.4% correct and 6.4% partially correct. In critical thinking tasks, ChatGPT-4 demonstrated higher accuracy, with 93.6% correct answers, compared to Gemini's 89.3% correct. For hypothetical scenarios, both models performed well, with ChatGPT-4 at 91.5% correct and Gemini at 87.2%. In open-ended discussions, ChatGPT-4 achieved 100% accuracy, whereas Gemini had 89.4% correct. However, both models showed room for improvement in conducting full comparisons. Further analysis of these partially correct responses will provide deeper insights into model performance and areas for improvement. Based on these results, the researchers argue for the strategic choice of fine-tuning ChatGPT-4 for a Q&A model tailored to supporting engineers in navigating the intricacies of the law and empower them with deeper insights into regulatory compliance, nuanced perspectives on applying legal mandates, and improved decision-making throughout their professional practice. This research marks a significant step in harnessing the power of AI for improved knowledge access and informed decision-making within the specific domain of engineering regulations in Bahrain. Further research recommends exploring domain-specific fine-tuning techniques and integration with collaborative platforms holds immense potential for unlocking the full potential of AI in streamlining compliance processes and enhancing overall efficiency within the engineering profession. Large Language Models (LLMs) can be leveraged to create a Q&A model for Engineering Regulations. One way to achieve this is by converting natural language (NL) requirements into machine-readable requirements using LLMs. This can be accomplished by creating a requirements table from free-form NL requirements and identifying boilerplate templates for different types of requirements based on linguistic patterns. By utilizing these language models together, the standardization of requirements can be achieved. To demonstrate the effectiveness of this approach, the researchers employed requirements from the law, with respect to regulating the practice of engineering professions in the Kingdom of Bahrain.