Dataset-centric AI ethics classification
摘要
AI ethics refers to the moral principles and guidelines governing the development and deployment of artificial intelligence systems, ensuring they align with human values and societal well-being. It encompasses the evaluation of AI outputs for fairness, safety, transparency, and respect for human rights. To advance systematic ethical evaluation, we introduce the EthicsLens dataset, comprising 38,808 responses generated by seven large language models. These responses were generated using diverse prompts designed to elicit appropriate and potentially sensitive responses. Each response was then annotated across sixteen ethical categories, including stereotyping, toxicity, misinformation, hate speech, harmful advice, privacy violations, political bias, false confidence, emotional or religious insensitivity, sexual content, manipulation, and impersonation. To classify ethical and unethical AI-generated content, the dataset is analysed using state-of-the-art classification methods, assessing its ability to support reliable ethical evaluation. Performance is reported both for binary ethical classification and multilabel violation identification. Results include accuracies of nearly 99% for binary classification tasks with SVM and CNN models, and macro-F1 scores of about 96% on multilabel tasks for Sentence-BERT transformer model.