Turing, A. M. (1950). Computing Machinery and
Intelligence.
Weizenbaum, J. (1966). ELIZA—A Computer Program for the Study of
Natural Language Communication Between Man and Machine.
Winograd, T. (1971). Procedures as a Representation for Data in
a Computer Program for Understanding Natural Language.
Weaver, W. (1949). Translation.
Weaver, W. (1952). Translation.
IBM (1954). 701 Translator / Machine-Aided Translation
相关材料.
ALPAC. (1966). Language and Machines: Computers in Translation
and Linguistics.
Brown, P. F., et al. (1990). A Statistical Approach to Machine
Translation.
Papineni, K., et al. (2002). BLEU: a Method for Automatic
Evaluation of Machine Translation.
Bahdanau, D., Cho, K., & Bengio, Y. (2014). Neural Machine
Translation by Jointly Learning to Align and Translate.
Bowman, S. R., et al. (2015). A Large Annotated Corpus for
Learning Natural Language Inference.
Rajpurkar, P., et al. (2016). SQuAD: 100,000+ Questions for
Machine Comprehension of Text.
Vaswani, A., et al. (2017). Attention Is All You Need.
Wang, A., et al. (2018). GLUE: A Multi-Task Benchmark and
Analysis Platform for Natural Language Understanding.
Devlin, J., et al. (2019). BERT: Pre-training of Deep
Bidirectional Transformers for Language Understanding.
Wang, A., et al. (2019). SuperGLUE: A Stickier Benchmark for
General-Purpose Language Understanding Systems.
Hu, J., et al. (2020). XTREME: A Massively Multilingual
Multi-task Benchmark for Evaluating Cross-lingual
Generalization.
Xu, L., et al. (2020). CLUE: A Chinese Language Understanding
Evaluation Benchmark.
Ouyang, L., et al. (2022). Training Language Models to Follow
Instructions with Human Feedback.
Yao, S., et al. (2022/2023). ReAct: Synergizing Reasoning and
Acting in Language Models.
Schick, T., et al. (2023). Toolformer: Language Models Can Teach
Themselves to Use Tools.
Bridgman, P. W. (1927). The Logic of Modern Physics.
Sapir, E. (1921). Language: An Introduction to the Study of
Speech.
Frege, G. (1892). On Sense and Reference.
Wittgenstein, L. (1953). Philosophical
Investigations(excerpt).
Yu, Z., et al. (2024). KIEval: A Knowledge-grounded Interactive
Evaluation Framework for Large Language Models.
链接:https://arxiv.org/abs/2402.15043
Grishman, R. and Sundheim, B. (1996). Message Understanding
Conference-6: A Brief History. In COLING 1996 Volume 1: The 16th
International Conference on Computational Linguistics, pages 466–471.
链接:https://aclanthology.org/C96-1079/
Weizenbaum, J. (1976). Computer Power and Human Reason: From
Judgment to Calculation. W. H. Freeman, San Francisco.
Zhu, K., et al. (2024). Dynamic Evaluation of Large Language
Models by Meta Probing Agents.
链接:https://arxiv.org/abs/2402.14865
Ni, R., et al. (2024). Benchmarking and Understanding
Compositional Relational Reasoning of LLMs.
链接:https://arxiv.org/abs/2412.12841
Cohen-Inger, N., et al. (2025). Forget What You Know about LLMs
Evaluations — LLMs are Like a Chameleon.
链接:https://arxiv.org/abs/2502.07445
Zhao, G., et al. (2025). Large Language Models Badly Generalize
across Option Length, Problem Types, and Irrelevant Noun
Replacements.
链接:https://arxiv.org/abs/2502.12459
Choukrani, O., et al. (2025). LLM-BABYBENCH: Understanding and
Evaluating Grounded Planning and Reasoning in LLMs.
链接:https://arxiv.org/abs/2505.12135