Context-sensitive Spelling Correction Using Google Web 1T 5-Gram Information

Computer Science – Computation and Language

Scientific paper

Rate now

  [ 0.00 ] – not rated yet Voters 0   Comments 0

Details

LACSC - Lebanese Association for Computational Sciences - http://www.lacsc.org

Scientific paper

10.5539/cis.v5n3p37

In computing, spell checking is the process of detecting and sometimes providing spelling suggestions for incorrectly spelled words in a text. Basically, a spell checker is a computer program that uses a dictionary of words to perform spell checking. The bigger the dictionary is, the higher is the error detection rate. The fact that spell checkers are based on regular dictionaries, they suffer from data sparseness problem as they cannot capture large vocabulary of words including proper names, domain-specific terms, technical jargons, special acronyms, and terminologies. As a result, they exhibit low error detection rate and often fail to catch major errors in the text. This paper proposes a new context-sensitive spelling correction method for detecting and correcting non-word and real-word errors in digital text documents. The approach hinges around data statistics from Google Web 1T 5-gram data set which consists of a big volume of n-gram word sequences, extracted from the World Wide Web. Fundamentally, the proposed method comprises an error detector that detects misspellings, a candidate spellings generator based on a character 2-gram model that generates correction suggestions, and an error corrector that performs contextual error correction. Experiments conducted on a set of text documents from different domains and containing misspellings, showed an outstanding spelling error correction rate and a drastic reduction of both non-word and real-word errors. In a further study, the proposed algorithm is to be parallelized so as to lower the computational cost of the error detection and correction processes.

No associations

LandOfFree

Say what you really think

Search LandOfFree.com for scientists and scientific papers. Rate them and share your experience with other people.

Rating

Context-sensitive Spelling Correction Using Google Web 1T 5-Gram Information does not yet have a rating. At this time, there are no reviews or comments for this scientific paper.

If you have personal experience with Context-sensitive Spelling Correction Using Google Web 1T 5-Gram Information, we encourage you to share that experience with our LandOfFree.com community. Your opinion is very important and Context-sensitive Spelling Correction Using Google Web 1T 5-Gram Information will most certainly appreciate the feedback.

Rate now

     

Profile ID: LFWR-SCP-O-313988

  Search
All data on this website is collected from public sources. Our data reflects the most accurate information available at the time of publication.