AI・テクノロジー研究

Linguistic Data Consortium

ldc.upenn.edu

AI・テクノロジー研究free + paid

言語研究やモデル訓練向けの音声・テキストコーパスを配布。利用は会員登録制だが、コーパスの説明ページは誰でも閲覧できる。

LDC has been the standard distributor of corpora used to train and benchmark NLP systems for decades — the Penn Treebank came through here — so its documentation pages describing how a corpus was collected and annotated carry real weight for a notebook about NLP methodology. Licensing a corpus itself requires paid membership, so LDC works better as a documentation and methodology source for a notebook than as a place to pull data files directly.

Linguistic Data Consortiumを開く

Gemini Notebook に取り込む

  1. 1上のリンク(またはサイト内の特定のページ)をコピーします。
  2. 2ノートブックで「ソースを追加」→「ウェブサイト」をクリックします。
  3. 3リンクを貼り付けて確定します。複数のリンクは1行に1件ずつ、まとめて貼り付けられます。

学会の予稿集や研究機関のブログでは、論文全文がPDFやWebページとして公開されています。これらはきれいに取り込め、引用情報もそのまま保たれます。

ai・テクノロジー研究の他のソース