This is a cross-sectional study that primarily employs quantitative analysis, supplemented by qualitative assessment. The research is conducted in two stages: Phase I consists of a model performance comparison experiment, and Phase II involves an item quality evaluation experiment. The entire study adheres to the principles of single-blinding, randomization, and standardization to ensure scientific rigor and reproducibility.
The single-blind design is implemented during the "standardized testing" phase, where the system intersperses AI-generated items with those authored by human experts. Participants remain blinded to the source of each item (AI-generated vs. human-authored) throughout the testing and scoring processes, thereby ensuring the objectivity of the evaluation results.