PENILAIAN EKSPERIMENTAL RALAT MAKLUMAT JANAAN KECERDASAN BUATAN BERASASKAN SET DATA KEFATWAAN
DOI:
https://doi.org/10.33102/jfatwa.vol31no3.821Keywords:
Fatwa, Ralat Maklumat Janaan AI, penjanaan diperkaya berasaskan capaian maklumat (RAG), sistem berasaskan pengetahuan, sistem maklumat Islam.Abstract
Penggunaan teknologi kecerdasan buatan (AI) semakin meluas dan berpotensi memperkukuh penyebaran maklumat Islam. Bagaimanapun, maklumat yang dijana oleh AI turut membawa risiko berlakunya ralat, khususnya apabila digunakan dalam bidang berautoriti seperti kefatwaan. Oleh itu, kajian eksploratori ini menilai ralat dalam maklumat janaan AI menggunakan set data berkaitan kefatwaan melalui reka bentuk eksperimen terkawal yang membandingkan tiga konfigurasi ChatGPT 5.2. Set data kajian distrukturkan kepada dua tema, iaitu Konsep Asas Fatwa dan Cabaran Institusi Fatwa. Respons AI dinilai berdasarkan empat dimensi ralat maklumat; iaitu konsistensi, koheren, keterangkuman dan halusinasi, dengan menggunakan rubrik berskala Likert lima mata dan skor median sebagai ukuran kecenderungan pusat bagi data berskala ordinal. Dalam skop kajian ini, konfigurasi penjanaan diperkaya berasaskan capaian maklumat tertutup menunjukkan kestabilan yang lebih baik berbanding dua konfigurasi lain dan berpotensi diteliti lebih lanjut sebagai satu tetapan dalam penggunaan AI untuk konteks kefatwaan. Hal ini berdasarkan skor median yang kekal pada tahap memuaskan, iaitu 4.25 bagi Tema 1 dan 4.00 bagi Tema 2, dengan perbezaan skor agregat tema sebanyak 0.25. Secara keseluruhan, dapatan kajian ini menggariskan keperluan kepada kerangka tadbir urus dan penilaian saintifik yang bersepadu sebelum teknologi AI boleh digunakan sebagai alat sokongan maklumat dalam ekosistem fatwa.
Downloads
References
AbuJarour, S., Qarariah, A., Saadeh, N., & Salem, M. (2024). AI, Misinformation, and Fake News: A Literature Review of Ethical and Technical Approaches. In Contributions to Finance and Accounting: Vol. Part F3769 (pp. 641–652). Springer Nature. https://doi.org/10.1007/978-3-031-67547-8_55
Al Kubaisi, A. A. S. H. (2024). Ethics of Artificial Intelligence a Purposeful and Foundational Study in Light of the Sunnah of Prophet Muhammad. Religions, 15(11). https://doi.org/10.3390/rel15111300
Aldoseri, A., Al-Khalifa, K. N., & Hamouda, A. M. (2023). Re-Thinking Data Strategy and Integration for Artificial Intelligence: Concepts, Opportunities, and Challenges. Applied Sciences (Switzerland), 13(12). https://doi.org/10.3390/app13127082
Alnefaie, S., Atwell, E., & Alsalka, M. A. (2025). Question Answering over the Arabic Hadith Sharif Using Transformer Models. Communications in Computer and Information Science, 2339 CCIS, 195–206. https://doi.org/10.1007/978-3-031-79164-2_17
Asgari, E., Montaña-Brown, N., Dubois, M., Khalil, S., Balloch, J., Yeung, J. A., & Pimenta, D. (2025). A framework to assess clinical safety and hallucination rates of LLMs for medical text summarisation. Npj Digital Medicine 2025 8:1, 8(1), 274. https://doi.org/10.1038/s41746-025-01670-7
Bdoor, S. Y., & Habes, M. (2025). Meta-Analysis of Ethical Challenges in the Use of Artificial Intelligence in Arab Media Institutions. Studies in Computational Intelligence, 1208, 131–146. https://doi.org/10.1007/978-3-031-89175-5_9
Brown, A., Roman, M., & Devereux, B. (2025). A Systematic Literature Review of Retrieval-Augmented Generation: Techniques, Metrics, and Challenges. ArXiv. https://doi.org/10.48550/arXiv.2508.06401
Bruno, A., Mazzeo, P. L., Chetouani, A., Tliba, M., & Kerkouri, M. A. (2023). Insights into Classifying and Mitigating LLMs’ Hallucinations. CEUR Workshop Proceedings. https://www.scopus.com/pages/publications/85178666482
Chiu, Edwin Kwan-Yeung, and Tom Wai-Hin Chung. (2025). Protocol for human evaluation of generative artificial intelligence chatbots in clinical consultations. PLoS ONE, 20(3). https://doi.org/10.1371/journal.pone.0300487
Dahl, M., Magesh, V., Suzgun, M., & Ho, D. E. (2024). Large Legal Fictions: Profiling Legal Hallucinations in Large Language Models. Journal of Legal Analysis, 16(1), 64–93. https://doi.org/10.1093/JLA/LAAE003
FACTS Team. (2024). FACTS Grounding: A new benchmark for evaluating the factuality of large language models - Google DeepMind. https://deepmind.google/blog/facts-grounding-a-new-benchmark-for-evaluating-the-factuality-of-large-language-models/
Fang, X., Yuan, H., Li, H., Kou, J., Gu, C., Zhang, W., Duan, X., & Fang, Y. (2024). Reframing Hallucination in Large Language Models: A Lifecycle-Based, Mechanism-Aligned, and Phenomenon-Consistent Definition. 7th International Conference on Universal Village, UV 2024. https://doi.org/10.1109/UV63228.2024.11189217
Farquhar, S., Kossen, J., Kuhn, L., & Gal, Y. (2024). Detecting hallucinations in large language models using semantic entropy. Nature 2024 630:8017, 630(8017), 625–630. https://doi.org/10.1038/s41586-024-07421-0
Gopali, S., Siami-Namini, S., Abri, F., & Namin, A. S. (2024). The performance of the LSTM-based code generated by Large Language Models (LLMs) in forecasting time series data. Natural Language Processing Journal, 9, 100120. https://doi.org/10.1016/J.NLP.2024.100120
Heo, S., Son, S., & Park, H. (2025). HaluCheck: Explainable and verifiable automation for detecting hallucinations in LLM responses. Expert Systems with Applications, 272, 126712. https://doi.org/10.1016/J.ESWA.2025.126712
Jacovi, A., Wang, A., Alberti, C., Tao, C., Lipovetz, J., Olszewska, K., Haas, L., Liu, M., Keating, N., Bloniarz, A., Saroufim, C., Fry, C., Marcus, D., Kukliansky, D., Singh Tomar, G., Swirhun, J., Xing, J., Wang, L., Gurumurthy, M., … Das, D. (2025). The FACTS Grounding Leaderboard: Benchmarking LLMs’ Ability to Ground Responses to Long-Form Input. arXiv preprint arXiv:2501.03200. https://arxiv.org/abs/2501.03200.
Jesson, A., Beltran-Velez, N., Chu, Q., Karlekar, S., & Kossen, J. (2024). Estimating the Hallucination Rate of Generative AI. Advances in Neural Information Processing Systems. https://www.scopus.com/pages/publications/105000550746
Kannike, U. M. M., & Fahm, A. O. (2025). Exploring The Ethical Governance of Artificial Intelligence From an Islamic Ethical Perspective. Jurnal Fiqh, 22(1), 134–161. https://doi.org/10.22452/fiqh.vol22no1.5
Kollar, J., & Alshibli, M. (2024). An Overview of Artificial Intelligence’s Accuracy. 2024 IEEE Long Island Systems, Applications and Technology Conference, LISAT 2024. https://doi.org/10.1109/LISAT63094.2024.10808042
Koutrika, G. (2025). Navigating the Challenges of AI-Driven Data Processing. DEBS 2025 - Proceedings of the 19th ACM International Conference on Distributed and Event-Based Systems, 1–2. https://doi.org/10.1145/3701717.3733846
Kulkarni, A., Zhang, Y., Moniz, J. R. A., Ge, X., Tseng, B.-H., Piraviperumal, D., & Swayamdipta, S. (2025). Evaluating evaluation metrics: The mirage of hallucination detection. ArXiv. https://arxiv.org/abs/2504.18114
Laskar, M. T. R., Alqahtani, S., Bari, M. S., Rahman, M., Khan, M. A. M., Khan, H., Jahan, I., Bhuiyan, M. A. H., Tan, C. W., Parvez, M. R., Hoque, E., Joty, S., & Huang, J. X. (2024). A Systematic Survey and Critical Review on Evaluating Large Language Models: Challenges, Limitations, and Recommendations. EMNLP 2024 - 2024 Conference on Empirical Methods in Natural Language Processing, Proceedings of the Conference, 13785–13816. https://doi.org/10.18653/v1/2024.emnlp-main.764
Linardon, J., Jarman, H. K., McClure, Z., Anderson, C., Liu, C., & Messer, M. (2025). Influence of Topic Familiarity and Prompt Specificity on Citation Fabrication in Mental Health Research Using Large Language Models: Experimental Study. JMIR Mental Health, 12(1), e80371. https://doi.org/10.2196/80371
Marey, A., Saad, A. M., Tanas, Y., Ghorab, H., Niemierko, J., Backer, H., & Umair, M. (2025). Evaluating the accuracy and reliability of AI chatbots in patient education. Egyptian Journal of Radiology and Nuclear Medicine, 56(1), 37. https://link.springer.com/content/pdf/10.1186/s43055-025-01452-x.pdf
Maylawati, D. S., Khosyi’ah, S., Rizqullah, N., Fajar, F. I., & Ramdhani, M. A. (2025). Chatbot Assistant for Enhancing Religious Court Services in Indonesia Using Deep Learning-based Algorithms. International Journal of Computing, 24(2), 243–253. https://doi.org/10.47839/IJC.24.2.4007
Meerangani, K. A., Harun, M. S., & Mansor, M. S. (2025). Prinsip Fiqh Dalam Penyebaran Maklumat Berasaskan Kecerdasan Buatan (AI) Di Media Sosial. Jurnal Fiqh, 22(2), 235–261. https://doi.org/10.22452/FIQH.VOL22NO2.2
Mortaheb, M., Khojastepour, M. A. A., Chakradhar, S. T., & Ulukus, S. (2025). Rag-check: Evaluating multimodal retrieval augmented generation performance. arXiv preprint arXiv:2501.03995.
Munshi, A. A., AlSabban, W. H., Farag, A. T., Rakha, O. E., Al Sallab, A. A., & Alotaibi, M. (2021). Towards an automated Islamic fatwa system: Survey, dataset and benchmarks. International Journal of Computer Science and Mobile Computing, 10(4). https://doi.org/10.47760/ijcsmc.2021.v10i04.017
Musleh Al-Sartawi, A. M. A. (Ed.). (2021). The Big Data-Driven Digital Economy: Artificial and Computational Intelligence (Vol. 974). Springer International Publishing. https://doi.org/10.1007/978-3-030-73057-4
Nascimento, N., Guimaraes, E., Chintakunta, S. S., & Boominathan, S. A. (2025). How Effective are LLMs for Data Science Coding? A Controlled Experiment. Proceedings - 2025 IEEE/ACM 22nd International Conference on Mining Software Repositories, MSR 2025, 211–222. https://doi.org/10.1109/MSR66628.2025.00041
Niri, M. A., Saadon, M. H. M., Nawawi, M. S. A. M., & Jamaludin, M. H. (2022). Integrasi Model DIKW (Data-Information-Knowledge-Wisdom) Dalam Ilmu Falak Berasaskan Kerangka Sains Islam. Afkar, 24(2), 99–142. https://doi.org/10.22452/afkar.vol24no2.3
Novikova, J., Anderson, C., Blili-Hamelin, B., Rosati, D., & Majumdar, S. (2025). Consistency in Language Models: Current Landscape, Challenges, and Future Directions. ArXiv. https://arxiv.org/pdf/2505.00268
Nurfikri, S. R., Handayani, D., Fahreza, A. M., Almadani, D., Suryady, Z., & Lathifah, Z. K. (2024). Fikr: AI Chatbot Powered with Vector Search. 2024 10th International Conference on Computing, Engineering and Design, ICCED 2024. https://doi.org/10.1109/ICCED64257.2024.10983626
Nurhayati, Abdurrohman, R., Hulliyah, K., & Khairani, D. (2024). Performance Evaluation of a Chatbot with Retrieval Augmented Generation and Generative Pre-trained Transformer-4 Model for Taharah Domain. 2024 9th International Conference on Informatics and Computing, ICIC 2024. https://doi.org/10.1109/ICIC64337.2024.10957519
Özdemir, Ö. T., Kavan, M. Y., & Güven, Y. (2025). Evaluation of the readability, quality, and accuracy of AI chatbots. BMC Oral Health, 25(1), 1812. https://link.springer.com/article/10.1186/s12903-025-07298-z
Raffinetti, E. (2023). A Rank Graduation Accuracy measure to mitigate Artificial Intelligence risks. Quality and Quantity, 57, 131–150. https://doi.org/10.1007/s11135-023-01613-y
Rahim, S. F. A., Rahman, M. F. A., Thaidi, H. A. A., Azimi, N. N. M. A., & Jailani, M. R. (2025). Artificial intelligence for fatwa issuance: Guidelines and ethical considerations. Journal of Fatwa Management and Research, 30(1). https://doi.org/10.33102/jfatwa.vol30no1.654
Rahman, N. N. A., & Niri, M. A. (2024). Cabaran institusi fatwa dalam era Revolusi Industri 4.0. In Isu-Isu Fiqh Revolusi Industri Keempat (pp. 24–38). Penerbit UTHM.
Rashed, A., & Shirmohammadi, S. (2022). A Novel Method to Estimate Measurement Error in AI-Assisted Measurements. Conference Record - IEEE Instrumentation and Measurement Technology Conference. https://doi.org/10.1109/I2MTC48687.2022.9806449
Santhoshkumar, S. P., Karthika, R., Vijayakumar, N., Gajalakshmi, S., Jayanthi, K., & Hariharasudhan, S. (2025). Analysis of Duplicating the Human Mind Suspicions in AI in Various Divisions. Proceedings of the International Conference on Multi-Agent Systems for Collaborative Intelligence, ICMSCI 2025, 1707–1714. https://doi.org/10.1109/ICMSCI62561.2025.10894681
Shiferaw, K. B. (2024). Assessing the accuracy and quality of artificial intelligence chatbot-generated information. BMC Medical Informatics and Decision Making, 24(1), 404. https://link.springer.com/article/10.1186/s12911-024-02824-5
Shuster, K., Poff, S., Chen, M., Kiela, D., & Weston, J. (2021). Retrieval Augmentation Reduces Hallucination in Conversation. ArXiv. https://arxiv.org/abs/2104.07567
Shweiki, D., Al-Ahmad, S., & Samhan, A. A. A. (2025). Using Artificial Intelligence in Issuing Fatwas: A Jurisprudential Study. Studies in Systems, Decision and Control, 572, 761–768. https://doi.org/10.1007/978-3-031-76011-2_53
Sun, Y., Sheng, D., Zhou, Z., & Wu, Y. (2024). AI hallucination: towards a comprehensive classification of distorted information in artificial intelligence-generated content. Humanities and Social Sciences Communications 2024 11:1, 11(1), 1278. https://doi.org/10.1057/s41599-024-03811-x
Vartumyan, A. A., Mikhailov, G. G., Galdin, E. V., Lavrova, T. N., & Orobinskaya, V. N. (2023). Risks Associated with the Use of Artificial Intelligence in Various Fields of Science. Proceedings - 2023 5th International Conference on Control Systems, Mathematical Modeling, Automation and Energy Efficiency, SUMMA 2023, 434–438. https://doi.org/10.1109/SUMMA60232.2023.10349424
Warrens, M. J. (2013). Cohen’s weighted kappa with additive weights. Advances in Data Analysis and Classification, 7(1), 41–55. https://doi.org/10.1007/s11634-013-0123-9
Zhang, W., & Zhang, J. (2025). Hallucination Mitigation for Retrieval-Augmented Large Language Models: A Review. Mathematics 2025, Vol. 13, Page 856, 13(5), 856. https://doi.org/10.3390/MATH13050856
Downloads
Published
Issue
Section
License
Copyright (c) 2026 Mohammaddin Abdul Niri, Noor Naemah Abdul Rahman, Mohd Anuar Ramli

This work is licensed under a Creative Commons Attribution 4.0 International License.
The copyright of this article will be vested to author(s) and granted the journal right of first publication with the work simultaneously licensed under the Creative Commons Attribution 4.0 International (CC BY 4.0) license, unless otherwise stated.











