# \[Feasibility\] Custom file format containing records and query it with prolog

**URL:** <https://swi-prolog.discourse.group/t/feasibility-custom-file-format-containing-records-and-query-it-with-prolog/2129>\
**Category:** Help!\
**Created:** [April 9, 2020, 12:15pm UTC](https://swi-prolog.discourse.group/t/feasibility-custom-file-format-containing-records-and-query-it-with-prolog/2129 "2020-04-09T12:15:01Z")\
**Posts on this page:** 11\
**Page:** 1

<div class="post-metadata">

**Author:** ![AntonioL](https://avatars.discourse-cdn.com/v4/letter/a/4491bb/32.png) [@AntonioL](https://swi-prolog.discourse.group/u/AntonioL)\
**Post date:** [April 9, 2020, 12:15pm UTC](https://swi-prolog.discourse.group/t/feasibility-custom-file-format-containing-records-and-query-it-with-prolog/2129/1 "2020-04-09T12:15:01Z")

</div>

I have some C++ code which reads/write this file format.  
This file format consists in a list of heterogeneous records.

Think of something like:

```prolog
Player(player_id=10, player_name=LeBron James)
Team(team_id=1, name=Lakers)
PlaysFor(player_id=10, team_id=1)

```

We have lot of infrastructure around this file format. So I was wondering whether it is feasible for me to do the following:  
I would like to write a Prolog predicate which is able to read this file format and have it generate facts so that I can leverage SWI-Prolog expressiveness, ideally I would like to accomplish this by writing a C++ extension. This seems in the realm of the possible but I cannot find any example.

Can you point me to an example or some relevant documentation for this purpose?

---

<div class="post-metadata">

**Author:** ![jan](https://yyz2.discourse-cdn.com/free1/user_avatar/swi-prolog.discourse.group/jan/32/4_2.png) [@jan](https://swi-prolog.discourse.group/u/jan)\
**Post date:** [April 9, 2020, 12:40pm UTC](https://swi-prolog.discourse.group/t/feasibility-custom-file-format-containing-records-and-query-it-with-prolog/2129/2 "2020-04-09T12:40:42Z")

</div>

If you already have C++ code reading this, adding a C++ predicate is probably the way to go.  
See [http://www.swi-prolog.org/pldoc/package/pl2cpp.html](http://www.swi-prolog.org/pldoc/package/pl2cpp.html)

---

<div class="post-metadata">

**Author:** ![AntonioL](https://avatars.discourse-cdn.com/v4/letter/a/4491bb/32.png) [@AntonioL](https://swi-prolog.discourse.group/u/AntonioL)\
**Post date:** [April 9, 2020, 12:44pm UTC](https://swi-prolog.discourse.group/t/feasibility-custom-file-format-containing-records-and-query-it-with-prolog/2129/3 "2020-04-09T12:44:09Z")

</div>

I had already given a quick look to that documentation.

What I do not understand is how can I add a new “Fact”.

let’s say that I create a Predicate, read\_from\_my\_fileformat with arity 1, taking in input the path to the file, how can I have add a “new fact” as I read the file?

Can you point me to the part of that documentation which you think is very relevant for my use case?

What I do not understand is how to yield a new fact from within the PREDICATE macro, basically.

I apologise if my wording is not _compliant_ and probably I am not using the standard prolog terminology!

---

<div class="post-metadata">

**Author:** ![jan](https://yyz2.discourse-cdn.com/free1/user_avatar/swi-prolog.discourse.group/jan/32/4_2.png) [@jan](https://swi-prolog.discourse.group/u/jan)\
**Post date:** [April 9, 2020, 12:50pm UTC](https://swi-prolog.discourse.group/t/feasibility-custom-file-format-containing-records-and-query-it-with-prolog/2129/4 "2020-04-09T12:50:00Z")

</div>

You have roughly two options. One is to define a predicate in C++ that reads the next fact from the file. Then you create a loop in Prolog that calls this predicate to get a new fact and than calls assertz/1 to add the fact to the Prolog database.

Another option is to do more or less the same from C++: you loop through the file and for every fact you find you can Prolog assert/z through the C++ interface. It all depends a bit on your expertise, what you already have in C++ and what you want in Prolog.

---

<div class="post-metadata">

**Author:** ![AntonioL](https://avatars.discourse-cdn.com/v4/letter/a/4491bb/32.png) [@AntonioL](https://swi-prolog.discourse.group/u/AntonioL)\
**Post date:** [April 9, 2020, 12:53pm UTC](https://swi-prolog.discourse.group/t/feasibility-custom-file-format-containing-records-and-query-it-with-prolog/2129/5 "2020-04-09T12:53:31Z")

</div>

I think that assert/z was the bit I was looking for!

I will now take my time to go through the documentation, I think your answers gave me all the material I needed to at least be able to come up with a prototype.

Thanks

---

<div class="post-metadata">

**Author:** ![AntonioL](https://avatars.discourse-cdn.com/v4/letter/a/4491bb/32.png) [@AntonioL](https://swi-prolog.discourse.group/u/AntonioL)\
**Post date:** [April 9, 2020, 12:55pm UTC](https://swi-prolog.discourse.group/t/feasibility-custom-file-format-containing-records-and-query-it-with-prolog/2129/6 "2020-04-09T12:55:02Z")

</div>

Is there anything else I need to consider if I went ahead with the assert/z call from C++ land?

Also, I am keeping an open mind about my use case, is the above something which makes sense from a prolog perspective?

My idea is to leverage prolog in order to add querying capability to my some external stuff.

---

<div class="post-metadata">

**Author:** ![EricGT](https://avatars.discourse-cdn.com/v4/letter/e/f1d935/32.png) [@EricGT](https://swi-prolog.discourse.group/u/EricGT)\
**Post date:** [April 9, 2020, 1:05pm UTC](https://swi-prolog.discourse.group/t/feasibility-custom-file-format-containing-records-and-query-it-with-prolog/2129/7 "2020-04-09T13:05:55Z")

</div>

> [@AntonioL](#):
>
> is the above something which makes sense from a prolog perspective?

Here is the [elevator pitch](https://en.wikipedia.org/wiki/Elevator_pitch).

Having done something similar but all on the Prolog side.

Reading a file is done with predicates like [read\_stream\_to\_codes/2](https://www.swi-prolog.org/pldoc/doc_for?object=read_stream_to_codes/2).

The parsing would be done with [DCG](https://www.swi-prolog.org/pldoc/man?section=DCG)

The creation of facts can be done with [library(persistency)](https://www.swi-prolog.org/pldoc/man?section=persistency)

The reading of large fact files quickly can be done with [Quick load files](https://www.swi-prolog.org/pldoc/man?section=qlf).

---

<div class="post-metadata">

**Author:** ![AntonioL](https://avatars.discourse-cdn.com/v4/letter/a/4491bb/32.png) [@AntonioL](https://swi-prolog.discourse.group/u/AntonioL)\
**Post date:** [April 9, 2020, 1:07pm UTC](https://swi-prolog.discourse.group/t/feasibility-custom-file-format-containing-records-and-query-it-with-prolog/2129/8 "2020-04-09T13:07:47Z")

</div>

Appreciate that, thanks to the both of you for your help in getting me started.

---

<div class="post-metadata">

**Author:** ![emacstheviking](https://yyz2.discourse-cdn.com/free1/user_avatar/swi-prolog.discourse.group/emacstheviking/32/228_2.png) [@emacstheviking](https://swi-prolog.discourse.group/u/emacstheviking)\
**Post date:** [April 9, 2020, 2:25pm UTC](https://swi-prolog.discourse.group/t/feasibility-custom-file-format-containing-records-and-query-it-with-prolog/2129/9 "2020-04-09T14:25:01Z")

</div>

Just a quick note @AntonioL abvout your source format:

```prolog
Player(player_id=10, player_name=LeBron James)
Team(team_id=1, name=Lakers)
PlaysFor(player_id=10, team_id=1)

```

If you replaced “(”, “=” and “,” with SPACE, then you could probably manage without a DCG as:

```prolog
Player player_id 10 player_name LeBron James
Team team_id 1 name Lakers
PlaysFor player_id 10 team_id 1

```

Then the line format is  
token1: “topic/category/record-type” whataver!  
token 2,3 (n,n+1) are just Key and Value

Don’t know if that helps much!

---

<div class="post-metadata">

**Author:** ![emacstheviking](https://yyz2.discourse-cdn.com/free1/user_avatar/swi-prolog.discourse.group/emacstheviking/32/228_2.png) [@emacstheviking](https://swi-prolog.discourse.group/u/emacstheviking)\
**Post date:** [April 9, 2020, 2:37pm UTC](https://swi-prolog.discourse.group/t/feasibility-custom-file-format-containing-records-and-query-it-with-prolog/2129/10 "2020-04-09T14:37:13Z")

</div>

```prolog
?- split_string("Player(player_id=10, player_name=LeBron James)", "(=,", ")", X), [Class|KV]=X.
X = ["Player", "player_id", "10", " player_name", "LeBron James"],
Class = "Player",
KV = ["player_id", "10", " player_name", "LeBron James"].

```

Gotta love SWI Prolog!

---

<div class="post-metadata">

**Author:** ![peter.ludemann](https://yyz2.discourse-cdn.com/free1/user_avatar/swi-prolog.discourse.group/peter.ludemann/32/48_2.png) [@peter.ludemann](https://swi-prolog.discourse.group/u/peter.ludemann)\
**Post date:** [April 9, 2020, 5:08pm UTC](https://swi-prolog.discourse.group/t/feasibility-custom-file-format-containing-records-and-query-it-with-prolog/2129/11 "2020-04-09T17:08:01Z")

</div>

I have a similar need, except that the input is in JSON. I output the the facts using `write_canonical` and then `consult` them in a separate program (a detail: I use [load\_files/2](https://www.swi-prolog.org/pldoc/man?predicate=load_files/2)).

One little trick is to create both a `.pl` and an empty `.qlf` file; the first time the file is consulted, the `.qlf` file is flagged as invalid and updated, so the next time you consult the file, it’s much faster (and if you update the `.pl` file, the `.qlf` file will also be automatically updated the next time the file is consulted).
