Convert zero-padded bytes to UTF-8 string

Question

Convert zero-padded bytes to UTF-8 string

29.4k views Asked by Matt Joiner At 22 February 2011 at 04:36

I'm unpacking several structs that contain 's' type fields from C. The fields contain zero-padded UTF-8 strings handled by strncpy in the C code (note this function's vestigial behaviour). If I decode the bytes I get a unicode string with lots of NUL characters on the end.

>>> b'hiya\0\0\0'.decode('utf8')
'hiya\x00\x00\x00'

I was under the impression that trailing zero bytes were part of UTF-8 and would be dropped automatically.

What's the proper way to drop the zero bytes?

Original Q&A

There are 3 answers

Adam Rosenfield On 22 February 2011 at 04:43

Use str.rstrip() to remove the trailing NULs:

>>> 'hiya\0\0\0'.rstrip('\0')
'hiya'

phobie On 19 January 2013 at 21:36

Unlike the split/partition-solution this does not copy several strings and might be faster for long bytearrays.

data = b'hiya\0\0\0'
i = data.find(b'\x00')
if i == -1:
  return data
return data[:i]

**Duncan** · Accepted Answer · 2011-02-22T09:02:52+00:00

Either rstrip or replace will only work if the string is padded out to the end of the buffer with nulls. In practice the buffer may not have been initialised to null to begin with so you might get something like b'hiya\0x\0'.

If you know categorically 100% that the C code starts with a null initialised buffer and never never re-uses it, then you might find rstrip to be simpler, otherwise I'd go for the slightly messier but much safer:

>>> b'hiya\0x\0'.split(b'\0',1)[0]
b'hiya'

which treats the first null as a terminator.

TechQA.

Convert zero-padded bytes to UTF-8 string

There are 3 answers

Related Questions in PYTHON

Related Questions in UNICODE

Related Questions in UTF-8

Related Questions in BYTE

Related Questions in STRNCPY

Popular Questions

Trending Questions