参考文献

  1. Turing, A. M. (1950). Computing Machinery and Intelligence.
  2. Weizenbaum, J. (1966). ELIZA—A Computer Program for the Study of Natural Language Communication Between Man and Machine.
  3. Winograd, T. (1971). Procedures as a Representation for Data in a Computer Program for Understanding Natural Language.
  4. Weaver, W. (1949). Translation.
  5. Weaver, W. (1952). Translation.
  6. IBM (1954). 701 Translator / Machine-Aided Translation 相关材料.
  7. ALPAC. (1966). Language and Machines: Computers in Translation and Linguistics.
  8. Brown, P. F., et al. (1990). A Statistical Approach to Machine Translation.
  9. Papineni, K., et al. (2002). BLEU: a Method for Automatic Evaluation of Machine Translation.
  10. Bahdanau, D., Cho, K., & Bengio, Y. (2014). Neural Machine Translation by Jointly Learning to Align and Translate.
  11. Bowman, S. R., et al. (2015). A Large Annotated Corpus for Learning Natural Language Inference.
  12. Rajpurkar, P., et al. (2016). SQuAD: 100,000+ Questions for Machine Comprehension of Text.
  13. Vaswani, A., et al. (2017). Attention Is All You Need.
  14. Wang, A., et al. (2018). GLUE: A Multi-Task Benchmark and Analysis Platform for Natural Language Understanding.
  15. Devlin, J., et al. (2019). BERT: Pre-training of Deep Bidirectional Transformers for Language Understanding.
  16. Wang, A., et al. (2019). SuperGLUE: A Stickier Benchmark for General-Purpose Language Understanding Systems.
  17. Hu, J., et al. (2020). XTREME: A Massively Multilingual Multi-task Benchmark for Evaluating Cross-lingual Generalization.
  18. Xu, L., et al. (2020). CLUE: A Chinese Language Understanding Evaluation Benchmark.
  19. Ouyang, L., et al. (2022). Training Language Models to Follow Instructions with Human Feedback.
  20. Yao, S., et al. (2022/2023). ReAct: Synergizing Reasoning and Acting in Language Models.
  21. Schick, T., et al. (2023). Toolformer: Language Models Can Teach Themselves to Use Tools.
  22. Bridgman, P. W. (1927). The Logic of Modern Physics.
  23. Sapir, E. (1921). Language: An Introduction to the Study of Speech.
  24. Frege, G. (1892). On Sense and Reference.
  25. Wittgenstein, L. (1953). Philosophical Investigations(excerpt).
  26. Yu, Z., et al. (2024). KIEval: A Knowledge-grounded Interactive Evaluation Framework for Large Language Models. 链接:https://arxiv.org/abs/2402.15043
  27. Grishman, R. and Sundheim, B. (1996). Message Understanding Conference-6: A Brief History. In COLING 1996 Volume 1: The 16th International Conference on Computational Linguistics, pages 466–471. 链接:https://aclanthology.org/C96-1079/
  28. Weizenbaum, J. (1976). Computer Power and Human Reason: From Judgment to Calculation. W. H. Freeman, San Francisco.
  29. Zhu, K., et al. (2024). Dynamic Evaluation of Large Language Models by Meta Probing Agents. 链接:https://arxiv.org/abs/2402.14865
  30. Ni, R., et al. (2024). Benchmarking and Understanding Compositional Relational Reasoning of LLMs. 链接:https://arxiv.org/abs/2412.12841
  31. Cohen-Inger, N., et al. (2025). Forget What You Know about LLMs Evaluations — LLMs are Like a Chameleon. 链接:https://arxiv.org/abs/2502.07445
  32. Zhao, G., et al. (2025). Large Language Models Badly Generalize across Option Length, Problem Types, and Irrelevant Noun Replacements. 链接:https://arxiv.org/abs/2502.12459
  33. Choukrani, O., et al. (2025). LLM-BABYBENCH: Understanding and Evaluating Grounded Planning and Reasoning in LLMs. 链接:https://arxiv.org/abs/2505.12135