parsing large compressed xml files, python

Question

parsing large compressed xml files, python

3.1k views Asked by Marcin At 03 December 2009 at 21:23

file  = BZ2File(SOME_FILE_PATH)
p = xml.parsers.expat.ParserCreate()
p.Parse(file)

Here's code that tries to parse xml file compressed with bz2. Unfortunately it fails with a message:

TypeError: Parse() argument 1 must be string or read-only buffer, not bz2.BZ2File

Is there a way to parse on the fly compressed bz2 xml files?

Note: p.Parse(file.read()) is not an option here. I want to parse a file which is larger than available memory, so I need to have a stream.

Original Q&A

There are 3 answers

Amber On 03 December 2009 at 21:28

Use .read() on the file object to read in the entire file as a string, and then pass that to Parse?

file  = BZ2File(SOME_FILE_PATH)
p = xml.parsers.expat.ParserCreate()
p.Parse(file.read())

Joe Koberg On 03 December 2009 at 21:42

Can you pass in an mmap()'ed file? That should take care of automatically paging the needed parts of the file in, and avoid memory overflow. Of course if expat builts a parse tree, it might still run out of memory.

http://docs.python.org/library/mmap.html

Memory-mapped file objects behave like both strings and like file objects. Unlike normal string objects, however, these are mutable. You can use mmap objects in most places where strings are expected; for example, you can use the re module to search through a memory-mapped file.

**Nick** · Accepted Answer · 2009-12-03T21:47:38+00:00

Nick On 03 December 2009 at 21:47 BEST ANSWER

Just use p.ParseFile(file) instead of p.Parse(file).

Parse() takes a string, ParseFile() takes a file handle, and reads the data in as required.

Ref: http://docs.python.org/library/pyexpat.html#xml.parsers.expat.xmlparser.ParseFile

TechQA.

parsing large compressed xml files, python

There are 3 answers

Related Questions in PYTHON

Related Questions in COMPRESSION

Related Questions in BZIP

Popular Questions

Trending Questions