Feedback
18 results found
  1. Webis-Sentences-17

    • webis.de
    • temir.org
    Published Feb 27, 2017
  2. Webis-Simple-Sentences-17 Corpus

    • zenodo.org
    • search.datacite.org
    Published Feb 27, 2017
  3. Webis-QSpell-17

    • webis.de
    • temir.org
    • +1more
    Published 2017
  4. Webis-Mnemonics-17

    • webis.de
    • temir.org
    • +1more
    Published 2017
  5. Positive and Negative Sentences

    • www.kaggle.com
    Updated Jan 19, 2018
  6. m

    Data for: Language Models, Surprisal and Fantasy in Slavic...

    • data.mendeley.com
    Updated Aug 29, 2018
  7. Bilingual English-Danish parallel corpus from The Danish Nature Agency...

    • data.europa.eu
    Updated Oct 10, 2019
  8. SNAP Memetracker

    • www.kaggle.com
    Updated Nov 21, 2016
  9. K

    Pre-trained Word Vectors for Spanish

    • www.kaggle.com
    Updated 9 ago. 2017
  10. Prison in India

    • data.world
    Updated Sep 11, 2018
  11. d

    Federal Justice Statistics Program Data Series

    • www.da-ra.de
    Published Mar 2, 1990
  12. g

    Data from: Improving Prison Classification Procedures in Vermont: Applying...

    • datasearch.gesis.org
    • www.da-ra.de
    Published Aug 5, 2015
  13. z

    A Canadian French Emotional Speech Dataset

    • zenodo.org
    Published Apr 17, 2018
  14. d

    Hazmat Routes (National)

    • catalog.data.gov
    • data.wu.ac.at
    Updated Aug 17, 2017
  15. d

    Data from: Public Support for Rehabilitation in Ohio, 1996

    • www.da-ra.de
    Published Feb 17, 1999
  16. d

    Data from: Evaluation of Boot Camps for Juvenile Offenders in Cleveland,...

    • www.da-ra.de
    Published Nov 2, 1999
  17. d

    National Jail Census Series

    • www.da-ra.de
    Published Jul 13, 1996
  18. E

    HF radar daily averaged surface currents from the MOOSE MEDTLN sites (Toulon...

    • erddap.osupytheas.fr
    Created Aug 23, 2018
  19. Not seeing a result you expected?
    Learn how you can add new datasets to our index.

Share
Facebook
Twitter
Email
Click to copy link
Link copied

Webis-Sentences-17

  • Dataset published Feb 27, 2017
Dataset provided by
Bauhaus University, Weimarhttp://www.uni-weimar.de/
The Web Technology & Information Systems Network
Authors
Stein, Benno; Lucks, Stefan; Kiesel, Johannes
Description

The Webis-Sentences-17 corpus is a collection of 3,369,618,811 sentences extracted from the ClueWeb12 web crawl. It is designed to allow for statistical analyses of human-written sentences. More details on the sentence extraction can be found in the associated publication. The Webis-Simple-Sentences-17 corpus contains 471,085,690 English sentences from the Webis-Sentences-17 corpus. The sentences were sampled to achieve a level of sentence complexity similar to the one of sentences that humans make up as a memory aid for remembering passwords. Sentence complexity was determined by syllables per word. Both corpora are split in training and test set as they are used in the associated publication. The test set is extracted from part 00 of the ClueWeb12, while the training set is extracted from the other parts.

Search
Clear search
Close search
Google apps
Main menu