You have developed a tool to do a natural-language requirements-engineering task, such as finding ambiguities in a requirements specification. This chapter explains with an extended example how to evaluate the effectiveness of the tool. The chapter describes a full evaluation of one tool for finding ambiguities in any natural-language requirements-engineering document, such as a requirements specification. Typically, the evaluator runs the tool on a natural-language document whose ambiguities are known to determine the tool’s recall and precision. The evaluator then decides from this recall and precision whether the tool is any good. Unfortunately, the evaluator has no basis on which to declare the tool’s recall and precision to be acceptable other than experience and judgment — a not very scientific basis. Making the evaluation more scientific requires comparing the tool’s recall and precision to those of humans doing the same task and doing the evaluation as a full-fledged experiment in which the certainty of the conclusions can be measured. The chapter explains with some scenarios involving trade-offs what other data might need to be gathered during the evaluation in order to put the evaluation on a more scientific basis. It then refers you to sources of more detailed explanations of general tool evaluation.

错误:搜索内容不能为空,请输入英文关键词
错误:关键词超出字数限制,请精简
高级检索

Empirical Evaluation of Tools for Hairy Natural Language Requirements Engineering Tasks

  • Daniel M. Berry

摘要

You have developed a tool to do a natural-language requirements-engineering task, such as finding ambiguities in a requirements specification. This chapter explains with an extended example how to evaluate the effectiveness of the tool. The chapter describes a full evaluation of one tool for finding ambiguities in any natural-language requirements-engineering document, such as a requirements specification. Typically, the evaluator runs the tool on a natural-language document whose ambiguities are known to determine the tool’s recall and precision. The evaluator then decides from this recall and precision whether the tool is any good. Unfortunately, the evaluator has no basis on which to declare the tool’s recall and precision to be acceptable other than experience and judgment — a not very scientific basis. Making the evaluation more scientific requires comparing the tool’s recall and precision to those of humans doing the same task and doing the evaluation as a full-fledged experiment in which the certainty of the conclusions can be measured. The chapter explains with some scenarios involving trade-offs what other data might need to be gathered during the evaluation in order to put the evaluation on a more scientific basis. It then refers you to sources of more detailed explanations of general tool evaluation.