# Read\_string/3 counts code points, right?

**URL:** <https://swi-prolog.discourse.group/t/read-string-3-counts-code-points-right/4948>\
**Category:** Help!\
**Created:** [January 27, 2022, 11:56pm UTC](https://swi-prolog.discourse.group/t/read-string-3-counts-code-points-right/4948 "2022-01-27T23:56:25Z")\
**Posts on this page:** 6\
**Page:** 1

<div class="post-metadata">

**Author:** ![ericzinda](https://avatars.discourse-cdn.com/v4/letter/e/b3f665/32.png) [@ericzinda](https://swi-prolog.discourse.group/u/ericzinda)\
**Post date:** [January 27, 2022, 11:56pm UTC](https://swi-prolog.discourse.group/t/read-string-3-counts-code-points-right/4948/1 "2022-01-27T23:56:25Z")

</div>

From read\_string/3:

> `read_string(+Stream, ?Length, -String)`
> 
> Read at most Length characters from Stream and return them in the string String. If Length is unbound, Stream is read to the end and Length is unified with the number of characters read.

I want to double check here since character encoding can get slippery.

The “… at most Length characters…” means “… at most Length Unicode code points…” that are read from `Stream`, right?

Edit: Assuming the answer is yes: Is there a predicate like `read_string` that will read N _bytes_ from a stream and convert to a string?

---

<div class="post-metadata">

**Author:** ![jan](https://yyz2.discourse-cdn.com/free1/user_avatar/swi-prolog.discourse.group/jan/32/4_2.png) [@jan](https://swi-prolog.discourse.group/u/jan)\
**Post date:** [January 28, 2022, 8:54am UTC](https://swi-prolog.discourse.group/t/read-string-3-counts-code-points-right/4948/2 "2022-01-28T08:54:52Z")

</div>

> [@ericzinda](#):
>
> The “… at most Length characters…” means “… at most Length Unicode code points…” that are read from `Stream` , right?

Yes.

> [@ericzinda](#):
>
> Edit: Assuming the answer is yes: Is there a predicate like `read_string` that will read N _bytes_ from a stream and convert to a string?

No 🙂 You can use stream\_range\_open/3. That creates a filter stream with a defined length. Next you can simply read up to end-of-file. So,

```prolog
    setup_call_cleanup(
         stream_range_open(Input, Tmp, [size(Bytes)]),
         ( set_steam(Tmp, encoding(utf8)),
           read_string(Tmp, _, String)
         ),
         close(Tmp)).

```

Now, if you need to read JSON, a Prolog term or something else you can directly read from a stream you do not need the intermediate string. This is used to deal with `Content-Length: Bytes` in HTTP streams.

---

<div class="post-metadata">

**Author:** ![Boris](https://yyz2.discourse-cdn.com/free1/user_avatar/swi-prolog.discourse.group/boris/32/7486_2.png) [@Boris](https://swi-prolog.discourse.group/u/Boris)\
**Post date:** [January 28, 2022, 9:28am UTC](https://swi-prolog.discourse.group/t/read-string-3-counts-code-points-right/4948/3 "2022-01-28T09:28:33Z")

</div>

> [@ericzinda](#):
>
> Is there a predicate like `read_string` that will read N _bytes_ from a stream and convert to a string?

Note that with Unicode in UTF-8, that predicate would have a bug if you provide the number of bytes and you don’t read the whole unicode character.

---

<div class="post-metadata">

**Author:** ![EricGT](https://avatars.discourse-cdn.com/v4/letter/e/f1d935/32.png) [@EricGT](https://swi-prolog.discourse.group/u/EricGT)\
**Post date:** [January 28, 2022, 9:43am UTC](https://swi-prolog.discourse.group/t/read-string-3-counts-code-points-right/4948/4 "2022-01-28T09:43:12Z")

</div>

> [@Boris](#):
>
> Note that with Unicode in UTF-8, that predicate would have a bug if you provide the number of bytes and you don’t read the whole Unicode character.

Can you clarify with an example where the bug occurs. I ask because this is my line of thought and you seem to have a different line of thought.

```prolog
Welcome to SWI-Prolog (threaded, 64 bits, version 8.5.5)
...

?- current_prolog_flag(encoding,E).
E = text.

?- open_string("→",Stream),read_string(Stream,L,String),string_codes(String,Codes).
Stream = <stream>(000000000655DEC0),
L = 1,
String = "→",
Codes = [8594].

?- set_prolog_flag(encoding,utf8).
true.

?- current_prolog_flag(encoding,E).
E = utf8.

?- open_string("→",Stream),read_string(Stream,L,String),string_codes(String,Codes).
Stream = <stream>(000000000655DEC0),
L = 1,
String = "→",
Codes = [8594].

```

Rightwards Arrow → ([ref](https://www.compart.com/en/unicode/U+2192))

When the length is set to 1, it still reads the Unicode character.

```prolog
?- open_string("→",Stream),read_string(Stream,1,String),string_codes(String,Codes).
Stream = <stream>(0000000006257D00),
String = "→",
Codes = [8594].

```

---

<div class="post-metadata">

**Author:** ![Boris](https://yyz2.discourse-cdn.com/free1/user_avatar/swi-prolog.discourse.group/boris/32/7486_2.png) [@Boris](https://swi-prolog.discourse.group/u/Boris)\
**Post date:** [January 28, 2022, 10:26am UTC](https://swi-prolog.discourse.group/t/read-string-3-counts-code-points-right/4948/5 "2022-01-28T10:26:11Z")

</div>

Without going too deep, a byte is 8 bits, so the largest integer you can represent with a byte is 256. In your example, the code 8594 cannot be a byte.

For the gory details, start reading here: [UTF-8 - Wikipedia](https://en.wikipedia.org/wiki/UTF-8#Encoding)

---

<div class="post-metadata">

**Author:** ![jamesnvc](https://yyz2.discourse-cdn.com/free1/user_avatar/swi-prolog.discourse.group/jamesnvc/32/14_2.png) [@jamesnvc](https://swi-prolog.discourse.group/u/jamesnvc)\
**Post date:** [January 28, 2022, 8:52pm UTC](https://swi-prolog.discourse.group/t/read-string-3-counts-code-points-right/4948/6 "2022-01-28T20:52:04Z")

</div>

```prolog
?- open_string("😀", Stream), 
   stream_range_open(Stream, Tmp, [size(1)]), 
   read_string(Tmp, 1, String), 
   close(Tmp), close(Stream).
Warning: <stream>(0x55eac48fbdb0):1:1: Illegal UTF-8 continuation
Stream = <stream>(0x55eac48fb620),
Tmp = <stream>(0x55eac48fbdb0),
String = "�".

```
