Semanlink - [1909.01066] Language Models as Knowledge Bases?

Tags:

Σχετικά με το έγγραφο αυτό

sl:arxiv_author :
sl:arxiv_firstAuthor : Fabio Petroni
sl:arxiv_num : 1909.01066
sl:arxiv_published : 2019-09-03T11:11:08Z
sl:arxiv_summary : Recent progress in pretraining language models on large textual corpora led to a surge of improvements for downstream NLP tasks. Whilst learning linguistic knowledge, these models may also be storing relational knowledge present in the training data, and may be able to answer queries structured as \"fill-in-the-blank\" cloze statements. Language models have many advantages over structured knowledge bases: they require no schema engineering, allow practitioners to query about an open class of relations, are easy to extend to more data, and require no human supervision to train. We present an in-depth analysis of the relational knowledge already present (without fine-tuning) in a wide range of state-of-the-art pretrained language models. We find that (i) without fine-tuning, BERT contains relational knowledge competitive with traditional NLP methods that have some access to oracle knowledge, (ii) BERT also does remarkably well on open-domain question answering against a supervised baseline, and (iii) certain types of factual knowledge are learned much more readily than others by standard language model pretraining approaches. The surprisingly strong ability of these models to recall factual knowledge without any fine-tuning demonstrates their potential as unsupervised open-domain QA systems. The code to reproduce our analysis is available at https://github.com/facebookresearch/LAMA.@en
sl:arxiv_title : Language Models as Knowledge Bases?@en
sl:arxiv_updated : 2019-09-04T09:33:20Z
sl:bookmarkOf : https://arxiv.org/abs/1909.01066
sl:creationDate : 2019-09-05
sl:creationTime : 2019-09-05T22:32:00Z

Πληροφορία αρχείου

Bookmark of: https://arxiv.org/abs/1909.01066

Linked From

How Much Knowledge Can You Pack Into the Parameters of a Language Model?

Tags:

> It has recently been observed that neural language
models trained on unstructured text can
implicitly store and retrieve knowledge using
natural language queries.

indeed, cf. Facebook's paper [Language Models as Knowledge Bases?](/doc/2019/09/_1909_01066_language_models_as)

> In this short paper,
we measure the practical utility of this
approach by fine-tuning pre-trained models to
answer questions without access to any external
context or knowledge.

> we show that a large language
model pre-trained on unstructured text can
attain competitive results on open-domain question
answering benchmarks without any access
to external knowledge

BUT:

>1. state-of-the-art results only with the largest model
which had 11 billion parameters.
>1. “open-book” models
typically provide some indication of what information
they accessed when answering a question
that provides a useful form of interpretability.
In contrast, our model distributes knowledge
in its parameters in an inexplicable way, which
precludes this form of interpretability.
>1. **the maximum-likelihood objective provides no guarantees as to whether
a model will learn a fact or not.**

So, what's the point? To be compared with this [IBM's paper](/doc/2019/09/_1909_04120_span_selection_pre): "a new pre-training task inspired by reading comprehension and an effort to avoid encoding general knowledge in the transformer network itself"

2020-02-11 About

Documents with similar tags (experimental)

[1909.07606] K-BERT: Enabling Language Representation with Knowledge Graph

Tags:

2020-03-08 About

[1912.01412] Deep Learning for Symbolic Mathematics

Tags:

2019-12-09 About