3 datasets found

W
Webis-QSeC-10
webis.de
anthology.aicmu.ac.cn
3256198
Updated 2010
+ more versions
Share
Facebook
Twitter
Email
Click to copy link
Link copied
Cite
Matthias Hagen; Martin Potthast; Benno Stein (2010). Webis-QSeC-10 [Dataset]. http://doi.org/10.5281/zenodo.3256198
Explore at:
3256198Available download formats
Unique identifier
https://doi.org/10.5281/zenodo.3256198
Dataset updated
2010
Dataset provided by
Bauhaus-Universität Weimar
The Web Technology & Information Systems Network
Friedrich Schiller University Jena
University of Kassel, hessian.AI, and ScaDS.AI
Authors
Matthias Hagen; Martin Potthast; Benno Stein
License
Attribution 4.0 (CC BY 4.0)https://creativecommons.org/licenses/by/4.0/
License information was derived automatically
Description
The Webis Query Segmentation Corpus 2010 (Webis-QSeC-10) contains segmentations for 53,437 web queries obtained from Mechanical Turk crowdsourcing (4,850 used for training in our CIKM 2012 paper). For each query, at least 10 MTurk workers were asked to segment the query. The corpus represents the distribution of their decisions.
Z
Webis Query Segmentation Corpus 2010 (Webis-QSeC-10)
data.niaid.nih.gov
live.european-language-grid.eu
+2more
Updated Jan 24, 2020
Share
Facebook
Twitter
Email
Click to copy link
Link copied
Cite
Potthast, Martin (2020). Webis Query Segmentation Corpus 2010 (Webis-QSeC-10) [Dataset]. https://data.niaid.nih.gov/resources?id=zenodo_3256197
Explore at:
Dataset updated
Jan 24, 2020
Dataset provided by
Potthast, Martin
Beyer, Anna
Hagen, Matthias
Stein, Benno
Bräutigam, Christof
License
Attribution 4.0 (CC BY 4.0)https://creativecommons.org/licenses/by/4.0/
License information was derived automatically
Description
The Webis Query Segmentation Corpus 2010 (Webis-QSeC-10) contains segmentations for 53,437 web queries obtained from Mechanical Turk crowdsourcing (4,850 used for training in our CIKM 2012 paper). For each query, at least 10 MTurk workers were asked to segment the query. The corpus represents the distribution of their decisions.

We provide the training and test sets as single folders in Zip archives containing several files. The files "...-queries.txt" contain the query strings and a unique ID for each query. The files "...-segmentations-crowdsourced.txt" contain the crowdsourced segmentations with their number of votes per query ID (see below for an example). The "data" folders contain all the data (n-gram frequencies, PMI values, POS tags, etc.) needed to replicate the evaluation results of our proposed segmentation algorithms. For convenience reasons, the folder "segmentations-of-algorithms" contain the segmentations that our proposed algorithms compute.

The original queries were extracted from the AOL query log, and range from 3 to 10 keywords in length. For each query at least 10 MTurk workers were asked to segment the query and their decisions are accumulated in the corpus. The examples below demonstrate two different cases.

Sample queries with internal ID (as in "Webis-QSeC-10-training-set-queries.txt"):

2315313155 harvard community credit union

1858084875 women's cycling tops

Sample segmentations (as in "webis-qsec-10-training-set-segmentations-crowdsourced.txt"):

2315313155 [(6, 'harvard community credit union'), (2, 'harvard community|credit union'), (1, 'harvard|community|credit union'), (1, 'harvard|community credit union')]

1858084875 [(5, "women's|cycling tops"), (2, "women's|cycling|tops"), (2, "women's cycling|tops"), (1, "women's cycling tops")]

Each query has a unique internal ID (e.g., 2315313155 in the first example) and the segmentations file contains at least 10 different decisions the MTurk workers made for that query. In the first example, 6 workers have all 4 keywords in one segment, 2 workers decided to break after the second word (denoted by a |) etc. Note that apostrophe in the second example (query ID 1858084875) is escaped by double quotes around the segmentation strings.
E
Webis Query Spelling Corpus 2017 (Webis-QSpell-17)
live.european-language-grid.eu
data.niaid.nih.gov
+1more
csv
Updated May 26, 2024
+ more versions
Share
Facebook
Twitter
Email
Click to copy link
Link copied
Cite
(2024). Webis Query Spelling Corpus 2017 (Webis-QSpell-17) [Dataset]. https://live.european-language-grid.eu/catalogue/corpus/7757
Explore at:
csvAvailable download formats
Dataset updated
May 26, 2024
License
Attribution 4.0 (CC BY 4.0)https://creativecommons.org/licenses/by/4.0/
License information was derived automatically
Description
The Webis Query Spelling Corpus 2017 (Webis-QSpell-17) contains 54,772 web queries that were manually spell-checked; for 9,171 queries alternative spelling variants are contained.As for segmentations of many of the queries (i.e., tagged concepts and phrases), please refer to the companion corpus Webis-QSeC-10.
Not seeing a result you expected?
Learn how you can add new datasets to our index.

Facebook

Twitter

Click to copy link

Link copied

Cite

Matthias Hagen; Martin Potthast; Benno Stein (2010). Webis-QSeC-10 [Dataset]. http://doi.org/10.5281/zenodo.3256198

Webis-QSeC-10

Explore at:

7 scholarly articles cite this dataset (View in Google Scholar)

3256198Available download formats

Unique identifier

https://doi.org/10.5281/zenodo.3256198

Dataset updated

2010

Dataset provided by

Bauhaus-Universität Weimar
The Web Technology & Information Systems Network
Friedrich Schiller University Jena
University of Kassel, hessian.AI, and ScaDS.AI

Authors

Matthias Hagen; Martin Potthast; Benno Stein

License

Attribution 4.0 (CC BY 4.0)https://creativecommons.org/licenses/by/4.0/
License information was derived automatically

Description

The Webis Query Segmentation Corpus 2010 (Webis-QSeC-10) contains segmentations for 53,437 web queries obtained from Mechanical Turk crowdsourcing (4,850 used for training in our CIKM 2012 paper). For each query, at least 10 MTurk workers were asked to segment the query. The corpus represents the distribution of their decisions.

Webis-QSeC-10

Webis Query Segmentation Corpus 2010 (Webis-QSeC-10)

Webis Query Spelling Corpus 2017 (Webis-QSpell-17)

Webis-QSeC-10See More Versions

Webis-QSeC-10