Researchers or developers need to evaluate AI systems' performance on understanding flight search intent.
May use self-built datasets or manual evaluation, which is costly and inconsistent.
Lack of standardized evaluation methods makes it hard to compare AI systems in flight search scenarios.