Saved datasets
Last updated
Download format
Croissant
Croissant is a format for Machine Learning datasets
Learn more about this at mlcommons.org/croissant.
Usage rights
License from data provider
Please review the applicable license to make sure your contemplated use is permitted.
Topic
Provider
Free
Cost to access
Described as free to access or have a license that allows redistribution.
4 datasets found
  1. W

    Webis-Sentences-17

    • webis.de
    205950
    Updated 2017
  2. Webis-Simple-Sentences-17 Corpus

    • zenodo.org
    application/gzip
    Updated Jan 24, 2020
  3. E

    Webis Abstractive Snippet Corpus 2020

    • live.european-language-grid.eu
    • zenodo.org
    json
    Updated Aug 19, 2023
    + more versions
  4. E

    Bilingual English-Norwegian parallel corpus from the National Contact Point...

    • live.european-language-grid.eu
    • data.europa.eu
    tmx
    Updated Dec 19, 2021
  5. Not seeing a result you expected?
    Learn how you can add new datasets to our index.

Share
FacebookFacebook
TwitterTwitter
Email
Click to copy link
Link copied
Close
Cite
Johannes Kiesel; Benno Stein; Stefan Lucks (2017). Webis-Sentences-17 [Dataset]. http://doi.org/10.5281/zenodo.205950

Webis-Sentences-17

Explore at:
2 scholarly articles cite this dataset (View in Google Scholar)
205950Available download formats
Dataset updated
2017
Dataset provided by
Bauhaus-Universität Weimar
The Web Technology & Information Systems Network
Authors
Johannes Kiesel; Benno Stein; Stefan Lucks
License

Attribution 4.0 (CC BY 4.0)https://creativecommons.org/licenses/by/4.0/
License information was derived automatically

Description

The Webis-Sentences-17 corpus is a collection of 3,369,618,811 sentences extracted from the ClueWeb12 web crawl. It is designed to allow for statistical analyses of human-written sentences. More details on the sentence extraction can be found in the associated publication.

Search
Clear search
Close search
Google apps
Main menu