Awesome Question Answering
Question Answeringを扱う資料や関連プロジェクトをまとめたAwesomeリストです。
目次
最近の動向
最近のQAモデル
- DilBert: Delaying Interaction Layers in Transformer-based Encoders for Efficient Open Domain Question Answering (2020)
- UnifiedQA: Crossing Format Boundaries With a Single QA System (2020)
- ProQA: Resource-efficient method for pretraining a dense corpus index for open-domain QA and IR. (2020)
- TYDI QA: A Benchmark for Information-Seeking Question Answering in Typologically Diverse Languages (2020)
- Retrospective Reader for Machine Reading Comprehension
- TANDA: Transfer and Adapt Pre-Trained Transformer Models for Answer Sentence Selection (AAAI 2020)
最近の言語モデル
- ELECTRA: Pre-training Text Encoders as Discriminators Rather Than Generators - ケビン・クラークら、ICLR、2020。
- TinyBERT: Distilling BERT for Natural Language Understanding - ショウキ・ジョーら、ICLR、2020。
- MINILM: Deep Self-Attention Distillation for Task-Agnostic Compression of Pre-Trained Transformers - ウェンホイ・ワンガら、arXiv、2020。
- T5: Exploring the Limits of Transfer Learning with a Unified Text-to-Text Transformer - コリント・ラフェルら、arXivプレプリント、2019。
- ERNIE: Enhanced Language Representation with Informative Entities - チェンヤン・チャングら、ACL、2019。
- XLNet: Generalized Autoregressive Pretraining for Language Understanding - チリン・ヤンら、arXivプレプリント、2019。
- ALBERT: A Lite BERT for Self-supervised Learning of Language Representations - 蘭真忠ら、arXiv予稿、2019年。
- RoBERTa: A Robustly Optimized BERT Pretraining Approach - 劉尹漢ら、arXiv予稿、2019年。
- DistilBERT, a distilled version of BERT: smaller, faster, cheaper and lighter - ヴァイクター・サンハら、arXiv、2019年。
- SpanBERT: Improving Pre-training by Representing and Predicting Spans - ジョシ・マンドルら、TACL、2019年。
- BERT: Pre-training of Deep Bidirectional Transformers for Language Understanding - デビリン・ジャクブら、NAACL 2019、2018年。
AAAI 2020
- TANDA: Transfer and Adapt Pre-Trained Transformer Models for Answer Sentence Selection - AAAI 2020のQA関連論文です。
ACL 2019
- Overview of the MEDIQA 2019 Shared Task on Textual Inference, Question Entailment and Question Answering, Asma Ben Abacha, et al., ACL-W 2019, Aug 2019.
- Towards Scalable and Reliable Capsule Networks for Challenging NLP Applications - 趙偉ら、ACL 2019、2019年6月。
- Cognitive Graph for Multi-Hop Reading Comprehension at Scale - 明丁ら、ACL 2019、6月2019。
- Real-Time Open-Domain Question Answering with Dense-Sparse Phrase Index - 申敏俊ら、ACL 2019、6月2019。
- Unsupervised Question Answering by Cloze Translation - パトリック・レーウィスら、ACL 2019、6月2019。
- SemEval-2019 Task 10: Math Question Answering - マーカス・ホプキンスら、ACL-W 2019、6月2019。
- Improving Question Answering over Incomplete KBs with Knowledge-Aware Reader - 熊文漢ら、ACL 2019、5月2019。
- Matching Article Pairs with Graphical Decomposition and Convolutions - 劉邦ら、ACL 2019、5月2019。
- Episodic Memory Reader: Learning what to Remember for Question Answering from Streaming Data - Moonsu Han 他、ACL 2019、2019年3月。
- Natural Questions: a Benchmark for Question Answering Research - Tom Kwiatkowski 他、TACL 2019、2019年1月。
- Textbook Question Answering with Multi-modal Context Graph Understanding and Self-supervised Open-set Comprehension - Daesik Kim 他、ACL 2019、2018年11月。
EMNLP-IJCNLP 2019
- Language Models as Knowledge Bases? - Fabio Petron 他、EMNLP-IJCNLP 2019、2019年9月。
- LXMERT: Learning Cross-Modality Encoder Representations from Transformers - Hao Tan 他、EMNLP-IJCNLP 2019、2019年12月。
- Answering Complex Open-domain Questions Through Iterative Query Generation - Peng Qi 他、EMNLP-IJCNLP 2019、2019年10月。
- KagNet: Knowledge-Aware Graph Networks for Commonsense Reasoning - リン ビル・ユチェン 他、EMNLP-IJCNLP 2019、9月2019。
- Mixture Content Selection for Diverse Sequence Generation - チョ・ジャエミン 他、EMNLP-IJCNLP 2019、9月2019。
- A Discrete Hard EM Approach for Weakly Supervised Question Answering - ミン・セウォン 他、EMNLP-IJCNLP 2019、9月2019。
Arxiv
- Investigating the Successes and Failures of BERT for Passage Re-Ranking - パディジェラ・ハーシス 他、arXivプレプリント、5月2019。
- BERT with History Answer Embedding for Conversational Question Answering - ク・チェン 他、arXivプレプリント、5月2019。
- Understanding the Behaviors of BERT in Ranking - キオ・ユファン 他、arXivプレプリント、4月2019。
- BERT Post-Training for Review Reading Comprehension and Aspect-based Sentiment Analysis - 、胡徐ら、arXiv予備論文、2019年4月。
- End-to-End Open-Domain Question Answering with BERTserini - 、楊偉ら、arXiv予備論文、2019年2月。
- A BERT Baseline for the Natural Questions - 、クリス・アルベーリら、arXiv予備論文、2019年1月。
- Passage Re-ranking with BERT - 、ロドリゴ・ノゲイラら、arXiv予備論文、2019年1月。
- SDNet: Contextualized Attention-based Deep Network for Conversational Question Answering - 、朱成光ら、arXiv、2018年12月。
データセット
- ELI5: Long Form Question Answering - QAデータセットに関する研究です。
- CODAH: An Adversarially-Authored Question Answering Dataset for Common Sense, Michael Chen, et al., RepEval 2019, Jun 2019.
QAについて
QAの種類
- 単一ターンQA:文脈を考慮せずに回答します
- 会話型QA:それまでの会話ターンを利用します
QAの下位分類
- 知識ベースQA
- 表・リストベースQA
- テキストベースQA
- コミュニティベースQA
- 視覚QA
QAシステムの前処理における分析と構文解析
言語分析
- Morphological analysis
- Named Entity Recognition(NER)
- 同音異義語・多義語分析
- 統語解析(依存構造解析)
- 意味認識
多くのQAシステムを構成する3要素
- 事実抽出
- 質問の理解
- 回答の生成
主な出来事
- Wolfram Alphaが2009年に回答エンジンを公開しました。
- IBM Watsonシステムが2011年に Jeopardy! のトップチャンピオンを破りました。
- AppleのSiriが2011年にWolfram Alphaの回答エンジンを統合しました。
- Googleは2012年にFreebase知識ベースを活用したKnowledge Graphを公開し、QAを導入しました。
- Amazon Echo/Alexa(2015年)、Google Home/Google Assistant(2016年)、INVOKE/MS Cortana(2017年)、HomePod(2017年)
システム
- IBM Watson - 、最先端の性能を達成。
- Facebook DrQA - SQuAD1.0データセットに適用されている。SQuAD2.0データセットはリリースされたが、DrQAはまだテストされていない。
- MIT media lab’s Knowledge graph - 自由に利用できる意味ネットワークであり、コンピュータが人々が使う語の意味を理解するのを助けるように設計されている。
QAコンペティション
| データセット | 言語 | 主催者 | 開始年 | 首位 | モデル | 状態 | 人間の性能超過 | |
|---|---|---|---|---|---|---|---|---|
| 0 | Story Cloze Test | English | Univ. of Rochester | 2016 | msap | Logistic regression | Closed | x |
| 1 | MS MARCO | English | Microsoft | 2016 | YUANFUDAO research NLP | MARS | Closed | o |
| 2 | MS MARCO V2 | English | Microsoft | 2018 | NTT Media Intelli. Lab. | Masque Q&A Style | Opened | x |
| 3 | SQuAD | English | Univ. of Stanford | 2018 | XLNet (single model) | XLNet Team | Closed | o |
| 4 | SQuAD 2.0 | English | Univ. of Stanford | 2018 | PINGAN Omni-Sinitic | ALBERT + DAAF + Verifier (ensemble) | Opened | o |
| 5 | TriviaQA | English | Univ. of Washington | 2017 | Ming Yan | - | Closed | - |
| 6 | decaNLP | English | Salesforce Research | 2018 | Salesforce Research | MQAN | Closed | x |
| 7 | DuReader Ver1. | Chinese | Baidu | 2015 | Tryer | T-Reader (single) | Closed | x |
| 8 | DuReader Ver2. | Chinese | Baidu | 2017 | renaissance | AliReader | Opened | - |
| 9 | KorQuAD | Korean | LG CNS AI Research | 2018 | Clova AI LaRva Team | LaRva-Kor-Large+ + CLaF (single) | Closed | o |
| 10 | KorQuAD 2.0 | Korean | LG CNS AI Research | 2019 | Kangwon National University | KNU-baseline(single model) | Opened | x |
| 11 | CoQA | English | Univ. of Stanford | 2018 | Zhuiyi Technology | RoBERTa + AT + KD (ensemble) | Opened | o |
出版物
- 論文
- “Learning to Skim Text”, Adams Wei Yu, Hongrae Lee, Quoc V. Le, 2017. : テキスト中の必要な部分だけを表示します
- “Deep Joint Entity Disambiguation with Local Neural Attention”, Octavian-Eugen Ganea and Thomas Hofmann, 2017.
- “BI-DIRECTIONAL ATTENTION FLOW FOR MACHINE COMPREHENSION”, Minjoon Seo, Aniruddha Kembhavi, Ali Farhadi, Hananneh Hajishirzi, ICLR, 2017.
- “Capturing Semantic Similarity for Entity Linking with Convolutional Neural Networks”, Matthew Francis-Landau, Greg Durrett and Dan Klei, NAACL-HLT 2016.
- “Entity Linking with a Knowledge Base: Issues, Techniques, and Solutions”, Wei Shen, Jianyong Wang, Jiawei Han, IEEE Transactions on Knowledge and Data Engineering(TKDE), 2014.
- “Introduction to “This is Watson”, IBM Journal of Research and Development, D. A. Ferrucci, 2012.
- “A survey on question answering technology from an information retrieval perspective”, Information Sciences, 2011.
- “Question Answering in Restricted Domains: An Overview”, Diego Mollá and José Luis Vicedo, Computational Linguistics, 2007
- “Natural language question answering: the view from here”, L Hirschman, R Gaizauskas, natural language engineering, 2001.
- エンティティ曖昧性解消/エンティティリンキング
コード
- BiDAF - 双方向注意フロー(BIDAF)ネットワークは、文の異なる粒度レベルにおける文脈を表現する多段階階層プロセスであり、早期の要約なしにクエリに意識を持つ文脈表現を獲得するための双方向注意フロー機構を用いている。
- Official; Tensorflow v1.2
- Paper
- QANet - 質問と回答のアーキテクチャは再帰ネットワークを必要としない。そのエンコーダは、すべてのコンボリューションと自己注意から構成されており、コンボリューションは局所的な相互作用をモデル化し、自己注意はグローバルな相互作用をモデル化する。
- Google; Unofficial; Tensorflow v1.5
- Paper
- R-Net - reading comprehension style question answering, which aims to answer questions from a given passage向けのAn end-to-end neural networks model。
- MS; Unofficially by HKUST; Tensorflow v1.5
- Paper
- R-Net-in-Keras - KerasでR-NETを再実装した。
- MS; Unofficial; Keras v2.0.6
- Paper
- DrQA - reading comprehension applied to open-domain question answering向けのDrQA is a system。
- Facebook; Official; Pytorch v0.4
- Paper
- BERT - Bidirectional Encoder Representations from Transformers. Unlike recent language representation models, BERT is designed to pre-train deep bidirectional representations by jointly conditioning on both left and right context in all layers向けのA new language representation model which stands。
- Google; Official implementation; Tensorflow v1.11.0
- Paper
講義
- Question Answering - Natural Language Processing - 質問応答を扱う講義です。
スライド
- Question Answering with Knowledge Bases, Web and Beyond - スコット・ウェンタウ・イとハオ・マ | ミシガン・リサーチ | 2016年
- Question Answering - ハソ・プラッターファン・インスティテュート マリアナ・ネヴス 医師 2017年
データセット集
データセット
- AI2 Science Questions v2.1(2017)
- Children’s Book Test
- データセットの概要です。
- CODAH Dataset
- DeepMind Q&A Dataset; CNN/Daily Mail
- データセットの概要です。
- 論文: https://arxiv.org/abs/1506.03340
- ELI5
- GraphQuestions
- データセットの概要です。
- LC-QuAD
- データセットの概要です。
- MS MARCO
- データセットの概要です。
- 論文: https://arxiv.org/abs/1611.09268
- MultiRC
- データセットの概要です。
- 論文: http://cogcomp.org/page/publication_view/833
- NarrativeQA
- データセットの概要です。
- 論文: https://arxiv.org/pdf/1712.07040v1.pdf
- NewsQA
- データセットの概要です。
- 論文: https://arxiv.org/pdf/1611.09830.pdf
- Qestion-Answer Dataset by CMU
- データセットの概要です。
- SQuAD1.0
- データセットの概要です。
- 論文: https://arxiv.org/abs/1606.05250
- SQuAD2.0
- データセットの概要です。
- 論文: https://arxiv.org/abs/1806.03822
- Story cloze test
- データセットの概要です。
- 論文: https://arxiv.org/abs/1604.01696
- TriviaQA
- データセットの概要です。
- 論文: https://arxiv.org/abs/1705.03551
- WikiQA
- データセットの概要です。
IBM Watson DeepQA研究チームの5年間の出版物
- 2015
- “Automated Problem List Generation from Electronic Medical Records in IBM Watson”, Murthy Devarakonda, Ching-Huei Tsou, IAAI, 2015.
- “Decision Making in IBM Watson Question Answering”, J. William Murdock, Ontology summit, 2015.
- “Unsupervised Entity-Relation Analysis in IBM Watson”, Aditya Kalyanpur, J William Murdock, ACS, 2015.
- “Commonsense Reasoning: An Event Calculus Based Approach”, E T Mueller, Morgan Kaufmann/Elsevier, 2015.
- 2014
- “Problem-oriented patient record summary: An early report on a Watson application”, M. Devarakonda, Dongyang Zhang, Ching-Huei Tsou, M. Bornea, Healthcom, 2014.
- “WatsonPaths: Scenario-based Question Answering and Inference over Unstructured Information”, Adam Lally, Sugato Bachi, Michael A. Barborak, David W. Buchanan, Jennifer Chu-Carroll, David A. Ferrucci*, Michael R. Glass, Aditya Kalyanpur, Erik T. Mueller, J. William Murdock, Siddharth Patwardhan, John M. Prager, Christopher A. Welty, IBM Research Report RC25489, 2014.
- “Medical Relation Extraction with Manifold Models”, Chang Wang and James Fan, ACL, 2014.
Microsoft Researchの5年間の出版物
- 2018
- “Characterizing and Supporting Question Answering in Human-to-Human Communication”, Xiao Yang, Ahmed Hassan Awadallah, Madian Khabsa, Wei Wang, Miaosen Wang, ACM SIGIR, 2018.
- “FigureQA: An Annotated Figure Dataset for Visual Reasoning”, Samira Ebrahimi Kahou, Vincent Michalski, Adam Atkinson, Akos Kadar, Adam Trischler, Yoshua Bengio, ICLR, 2018
- 2017
- “Multi-level Attention Networks for Visual Question Answering”, Dongfei Yu, Jianlong Fu, Tao Mei, Yong Rui, CVPR, 2017.
- “A Joint Model for Question Answering and Question Generation”, Tong Wang, Xingdi (Eric) Yuan, Adam Trischler, ICML, 2017.
- “Two-Stage Synthesis Networks for Transfer Learning in Machine Comprehension”, David Golub, Po-Sen Huang, Xiaodong He, Li Deng, EMNLP, 2017.
- “Question-Answering with Grammatically-Interpretable Representations”, Hamid Palangi, Paul Smolensky, Xiaodong He, Li Deng,
- “Search-based Neural Structured Learning for Sequential Question Answering”, Mohit Iyyer, Wen-tau Yih, Ming-Wei Chang, ACL, 2017.
- 2016
- “Stacked Attention Networks for Image Question Answering”, Zichao Yang, Xiaodong He, Jianfeng Gao, Li Deng, Alex Smola, CVPR, 2016.
- “Question Answering with Knowledge Base, Web and Beyond”, Yih, Scott Wen-tau and Ma, Hao, ACM SIGIR, 2016.
- “NewsQA: A Machine Comprehension Dataset”, Adam Trischler, Tong Wang, Xingdi Yuan, Justin Harris, Alessandro Sordoni, Philip Bachman, Kaheer Suleman, RepL4NLP, 2016.
- “Table Cell Search for Question Answering”, Sun, Huan and Ma, Hao and He, Xiaodong and Yih, Wen-tau and Su, Yu and Yan, Xifeng, WWW, 2016.
- 2015
- “WIKIQA: A Challenge Dataset for Open-Domain Question Answering”, Yi Yang, Wen-tau Yih, and Christopher Meek, EMNLP, 2015.
- “Web-based Question Answering: Revisiting AskMSR”, Chen-Tse Tsai, Wen-tau Yih, and Christopher J.C. Burges, MSR-TR, 2015.
- “Open Domain Question Answering via Semantic Enrichment”, Huan Sun, Hao Ma, Wen-tau Yih, Chen-Tse Tsai, Jingjing Liu, and Ming-Wei Chang, WWW, 2015.
- 2014
- “An Overview of Microsoft Deep QA System on Stanford WebQuestions Benchmark”, Zhenghao Wang, Shengquan Yan, Huaming Wang, and Xuedong Huang, MSR-TR, 2014.
- “Semantic Parsing for Single-Relation Question Answering”, Wen-tau Yih, Xiaodong He, Christopher Meek, ACL, 2014.
Google AIの5年間の出版物
- 2018
- Google QA
- “QANet: Combining Local Convolution with Global Self-Attention for Reading Comprehension”, Adams Wei Yu, David Dohan, Minh-Thang Luong, Rui Zhao, Kai Chen, Mohammad Norouzi, Quoc V. Le, ICLR, 2018.
- “Ask the Right Questions: Active Question Reformulation with Reinforcement Learning”, Christian Buck and Jannis Bulian and Massimiliano Ciaramita and Wojciech Paweł Gajewski and Andrea Gesmundo and Neil Houlsby and Wei Wang, ICLR, 2018.
- “Building Large Machine Reading-Comprehension Datasets using Paragraph Vectors”, Radu Soricut, Nan Ding, 2018.
- Sentence representation
- “An efficient framework for learning sentence representations”, Lajanugen Logeswaran, Honglak Lee, ICLR, 2018.
- “Did the model understand the question?”, Pramod K. Mudrakarta and Ankur Taly and Mukund Sundararajan and Kedar Dhamdhere, ACL, 2018.
- Google QA
- 2017
- “Analyzing Language Learned by an Active Question Answering Agent”, Christian Buck and Jannis Bulian and Massimiliano Ciaramita and Wojciech Gajewski and Andrea Gesmundo and Neil Houlsby and Wei Wang, NIPS, 2017.
- “Learning Recurrent Span Representations for Extractive Question Answering”, Kenton Lee and Shimi Salant and Tom Kwiatkowski and Ankur Parikh and Dipanjan Das and Jonathan Berant, ICLR, 2017.
- Identify the same question
- “Neural Paraphrase Identification of Questions with Noisy Pretraining”, Gaurav Singh Tomar and Thyago Duque and Oscar Täckström and Jakob Uszkoreit and Dipanjan Das, SCLeM, 2017.
- 2014
- “Great Question! Question Quality in Community Q&A”, Sujith Ravi and Bo Pang and Vibhor Rastogi and Ravi Kumar, ICWSM, 2014.
Facebook AI Researchの5年間の出版物
- 2018
- Embodied Question Answering, Abhishek Das, Samyak Datta, Georgia Gkioxari, Stefan Lee, Devi Parikh, and Dhruv Batra, CVPR, 2018
- Do explanations make VQA models more predictable to a human?, Arjun Chandrasekaran, Viraj Prabhu, Deshraj Yadav, Prithvijit Chattopadhyay, and Devi Parikh, EMNLP, 2018
- Neural Compositional Denotational Semantics for Question Answering, Nitish Gupta, Mike Lewis, EMNLP, 2018
- 2017
- DrQA
- Reading Wikipedia to Answer Open-Domain Questions, Danqi Chen, Adam Fisch, Jason Weston & Antoine Bordes, ACL, 2017.
- DrQA
書籍
- Natural Language Question Answering system Paperback - Boris Galitsky (2003)
- New Directions in Question Answering - Mark T. Maybury (2004)
- Part 3. 5. Question Answering in The Oxford Handbook of Computational Linguistics - Sanda Harabagiu and Dan Moldovan (2005)
- Chap.28 Question Answering in Speech and Language Processing - Daniel Jurafsky & James H. Martin (2017)
リンク
- Building a Question-Answering System from Scratch— Part 1
- Qeustion Answering with Tensorflow By Steven Hewitt, O’REILLY, 2017
- Why question answering is hard
コントリビューション
コントリビューションを歓迎します。最初にコントリビューションガイドラインをお読みください。
ライセンス
法律で認められる限り、メンテナーの seriousmac は本作品に関するすべての著作権および関連・隣接権を放棄しています。