Skip to main content
Have a personal or library account? Click to login
Google Books Ngrams Recompressed and Searchable Cover
Open Access
|Dec 2012

References

  1. [1] Brants T., Popat A. C., Xu P., Och F. J., Dean J., Large language models in machine translation, in:, Prague, ACL 2007, 858-867.
  2. [2] Gao J., Nguyen P., Li X., Thrasher C., Li M., Wang K., A Comparative Study of Bing Web N-gram Language Models for Web Search and Natural Language Processing, in:, Geneva 2010.
  3. [3] Grabowski Sz., Swacha J., Compact Representation of URL Collections with Fast Access,,, 3, 2011, 349-355.
  4. [4] Guthrie D., Hepple M., Liu W., Efficient Minimal Perfect Hash Language Models, in: N. Calzolari, K. Choukri, B. Maegaard, J. Mariani, J. Odijk, S. Piperidis, M. Rosner, D. Tapias (eds.),, Valetta, ELRA 2010.
  5. [5] Michel J.-B. B., Kui Y., Presser A., Veres A., Gray M. K., Google Books Team, Picket J. P., Hoiberg D., Clancy D., Norvig P., Orwant J., Pinker S., Nowak M. A., Lieberman Aider E., Quantitative Analysis of Culture Using Millions of Digitized Books,,, 6014, 2011, 176-182.
  6. [6] Microsoft Research, Spelling Alteration for Web Search Workshop, City Center - Bellevue, WA, July 19, 2011. Materials available at http://webngram. research.microsoft.com/Spellerchallenge/Docs/Spelling_Alteration_Workshop. pdf (last checked: June 2012).
  7. [7] Pauls A., Klein D., Faster and Smaller N-Gram Language Models, in: Y. Matsumoto, R. Mihalcea (eds.),, Stroudsburg, ACL 2011, 258-267.
  8. [8] Procházka V., Pollák P., Analysis of Czech Web 1T 5-Gram Corpus and Its Comparison with Czech National Corpus Data, in: P. Sojka, A. Horák, I. Kopecek, K. Pala (eds.),, Brno, Springer 2010, 181-188.
  9. [9] Skibiński P., Grabowski Sz., Swacha J., Effective asymmetric XML compression,,, 10, 2008, 1027-1047.
  10. [10] Talbot D., Brants T., Randomized Language Models via Perfect Hash Functions, in:, Columbus, ACL 2008, 505-513.
  11. [11] Witten I. H., Moffat A., Bell T. C.,, Morgan Kaufmann Publishers, Los Altos, 1999.
  12. [12] Ziv, J., Lempel, A., A Universal Algorithm for Sequential Data Compression,,, 3, 1977, 337-343.
  13. [13] http://books.google.com/ngrams (last checked: June 2012).
  14. [14] http://books.google.com/ngrams/datasets (last checked: June 2012).
  15. [15] http://books.google.com/ngrams/info (last checked: June 2012).
  16. [16] http://iiwz.wneiz.pl/jakubs/progs/ngram_compressor.zip (last checked: June 2012).
  17. [17] http://research.microsoft.com/en-us/collaboration/focus/cs/web-ngram.aspx (last checked: June 2012).
  18. [18] http://www.base2ti.com (last checked: June 2012).
  19. [19] http://www.ldc.upenn.edu/Catalog/CatalogEntry.jsp?catalogId=LDC2006T13 (last checked: June 2012).
  20. [20] http://www.ldc.upenn.edu/Catalog/catalogEntry.jsp?catalogId=LDC2011T07 (last checked: June 2012).
DOI: https://doi.org/10.2478/v10209-011-0015-8 | Journal eISSN: 2300-3405 (formerly 0867-6356) | Journal ISSN: 0867-6356
Language: English
Page range: 271 - 281
Published on: Dec 22, 2012
Published by: Poznan University of Technology
In partnership with: Paradigm Publishing Services

© 2012 Szymon Grabowski, Jakub Swacha, published by Poznan University of Technology
This work is licensed under the Creative Commons License.