Python (nltk) - UnicodeDecodeError: 'ascii' codec can't decode byte

Question

I'm new to NLTK. I'm getting this error and I've searched around for encoding/decoding and specifically the UnicodeDecodeError but this error seems specific to the NLTK source code.

Here's the error:

Traceback (most recent call last):
  File "A:\Python\Projects\Test\main.py", line 2, in <module>
    print(pos_tag(word_tokenize("John's big idea isn't all that bad.")))
  File "A:\Python\Python\lib\site-packages\nltk\tag\__init__.py", line 100, in pos_tag
    tagger = load(_POS_TAGGER)
  File "A:\Python\Python\lib\site-packages\nltk\data.py", line 779, in load
    resource_val = pickle.load(opened_resource)
UnicodeDecodeError: 'ascii' codec can't decode byte 0xcb in position 0: ordinal not in range(128)

How do I go around fixing this error?

Here's what causes the error:

from nltk import pos_tag, word_tokenize
print(pos_tag(word_tokenize("John's big idea isn't all that bad.")))

Answer 1

try this... NLTK 3.0.1 with Python 2.7.x

import io
f = io.open(txtFile, 'rU', encoding='utf-8')

Answer 2

I had the same problem with you. I use Python 3.4 in Windows 7.

I had installed the "nltk-3.0.0.win32.exe" (from here ). But when i installed the "nltk-3.0a4.win32.exe" (from here ), my problem with nltk.pos_tag was solved. Check it.

EDIT: If the second link doesn't work, you can look here .

Answer 3

Duplicate: NLTK 3 POS_TAG throws UnicodeDecodeError

Long story short: NLTK isn't compatible with Python 3. You have to use NLTK 3 which sounds a bit experimental at this point.

Answer 4

Try using the module "textclean"

>>> pip install textclean

Python code

from textclean.textclean import textclean
text = textclean.clean("John's big idea isn't all that bad.")
print pos_tag(word_tokenize(text))

Python (nltk) - UnicodeDecodeError: 'ascii' codec can't decode byte

Question

4 answers

solution1
5 2015-01-14 17:06:43

solution2
4 2014-09-28 17:39:39

solution3
-2 2014-09-03 02:22:38

solution4
-2 2014-09-05 09:36:40

Python (nltk) - UnicodeDecodeError: 'ascii' codec can't decode byte

Question

4 answers

solution1 5 2015-01-14 17:06:43

solution2 4 2014-09-28 17:39:39

solution3 -2 2014-09-03 02:22:38

solution4 -2 2014-09-05 09:36:40

solution1
5 2015-01-14 17:06:43

solution2
4 2014-09-28 17:39:39

solution3
-2 2014-09-03 02:22:38

solution4
-2 2014-09-05 09:36:40