Вход на сайт

Просмотр новости

Найдите то, что Вас интересует

13th Web-as-Corpus Workshop @ EMNLP 2026

Дата публикации: 29-06-2026 00:00:00

The WaC-13 workshop invites research submissions on web data, corpus building, and linguistic analysis.

Основное содержимое страницы с новостью.

Felső-Víziváros and Matthias Church seen from the Danube

Felső-Víziváros és Mátyás-templom látványa a Dunáról · Thaler Tamas, Wikimedia Commons, CC BY-SA 3.0

On 29th October 2026, the Web-as-Corpus Workshop will return for its 13th edition, co-located with EMNLP 2026 in Budapest. The organising committee includes engineers from the Common Crawl Foundation alongside researchers from the Jožef Stefan Institute, the University of Oslo, and the University of Turku.

Research on the web as a corpus can be split into two main strands: using it as core data infrastructure for modern natural language processing, including Large Language Models (LLMs), or studying it as an object of societal and linguistic analysis in its own right. In both cases, questions about the curation, analysis, and responsible use of web-derived data have become increasingly critical. For those building systems, the "more is better" paradigm is under pressure from machine-generated content, data toxicity, limited metadata, and sparse data for many languages and domains. For those studying the web, issues around quality, representativeness, and ethical implications of data use are increasingly consequential. Both strands depend on understanding web data well enough to use it responsibly.

The WaC-13 workshop aims to connect researchers from multiple disciplines who share an interest in the web as a corpus. We invite submissions on methods, resources, and applications related to web corpora, with special emphasis on multilingual data and less-resourced languages.

Topics of interest include (but are not limited to):

  • Creation and evaluation of high-quality datasets for foundation models (e.g., data collection, filtering, enrichment, language identification)
  • Use of web data in empirical linguistic research
  • Analysis of web-scale corpora for quality, representativeness, and societal insights
  • Ethical and legal aspects of collecting, sharing, and using web data

There are two ways to submit your research: either directly by 7 August 2026, or through committing a pre-reviewed paper via ACL Rolling Review by 1 September 2026 (both deadlines AoE). Full details are available on the workshop website: https://wackyworkshop.org.

By bringing together researchers from NLP, linguistics, and the social sciences, WaC-13 aims to advance best practices for one of the field’s most influential data sources. If you work on any part of how web data is built, filtered, or studied, we’d love to read your submission.

Схожие новости

#Наименование новостиТональностьИнформативностьДата публикации
1Common Crawl Foundation at IIPC-WAC 2026017.610-06-2026
2קריקטורה יומית | 13 באוגוסט 202605012-08-2026
3קריקטורה יומית | 13 ביולי 20260012-07-2026
4CommonLID Update: New Tools, Growing Impact013.3516-06-2026
512 Aug 2026 13:00 : Learning to See, Generate, and Act for Scalable Robotic Manipulation031.4306-08-2026
612 Aug 2026 13:00 : Learning to See, Generate, and Act for Scalable Robotic Manipulation031.4306-08-2026
72026 ACM Conferentie over Reproducibility and Replicability0020-07-2026
8Save the date: Computational Research Symposium 2026019.4113-03-2026
912 Aug 2026 14:00 : Learning to See, Generate, and Act for Scalable Robotic Manipulation031.4306-08-2026
1028 Jul 2026 12:00 : Cognitive Memory Mechanisms for Understanding and Improving Large Language Models031.4317-07-2026

Классификация: Пресс-релизы. Схожих патентов: 0. Схожих новостей: 10. Тональность: 0. Информативность: 18.95. Источник: commoncrawl.org.