Hacker Newsnew | past | comments | ask | show | jobs | submitlogin

How do filters work? They seem pretty difficult implementation-wise since you can write them in any of the language bindings. My first guess is that you pipe all the data in a table to the client, and the client itself does the filtration. But this would be extraordinarily inefficient.


Piping all the data to the client would be extremely inefficient. Fortunately we don't do that.

When a filter is written in the client language it gets compiled into a protocol buffer which is sent to the cluster. This gets compiled into a query which is sent to each of the relevant shards for the table. This query has the filter baked right into it. The shards then go through their local copy of the data and filter out the rows which do not meet the query predicate. This data gets returned to the coordinating node and eventually to the user. Thus only the data the will actually be returned is ever transferred over the network.

Furthermore this process is done lazily. On the client side rather than getting back a huge array with the results of your filter you get back an iterator. This iterator stores a buffer of data which will be refilled as it is incremented.


To add to jdoliner's answer, the reason why you can write table('foo').filter(lambda x: x['bar'] > 5).run() is because we do some language trickery on the client side to compile the query to an AST. In this case, we overload greater than operator, call the lambda function once on the client with a special object, and return an AST. This AST is then sent to the server and executed there.

It's rather difficult to integrate into a host language like that smoothly from the driver implementation perspective, but once the driver is written the user experience is amazing because you can write queries that look exactly like Python, but they're executed entirely on the server.


I find I prefer SQLAlchemy's Table.query.filter(Table.bar > 5) to a lambda that gets compiled to an AST in an odd way.


You can do that too: r.table('foo').filter(r['bar'] > 5)

The use of r in filter is getting the attribute bar of the row.


That's pretty cool.

Have you found that there are useful expressions that are awkward to express without the lambda trick?


Lambda is necessary because when you do nested subqueries, saying r['x'] is ambiguous and can cause all sorts of unpleasantness. So, if you use nested queries, the server rejects the implicit syntax and requires the use of lambda.

Lambda syntax is really nice too, I actually prefer it for writing queries.


I find the lambda trick not explicit and obvious enough. I fear I would do something stupid like trigger a side-effect without realising.

Again taking an example from SQLAlchemy, you can explicitly make a subquery, and then reference it instead of the original Table. A binding more like SQLAlchemy can probably written for RethinkDB.


Does the function have to be referentially transparent?


No. E.g., you could write r.table('foo').update(lambda row: row.merge({'bar': row['bar'] + 1 })). A shortcut for this is r.table('foo').update({'bar': r['bar'] + 1 }). Neither is referentially transparent, and both work.


I believe both of those functions are referential transparent because they're pure functions. An example of a non referential transparent function would be:

r.table('foo').update(lambda row: {'bar' : r.table('bar').get(row["bar_id"]))

This still works but gets evaluated in a different way to make sure every secondary winds up with the same value.


This is beyond awesome, and thanks for the follow-ups!


Nope, they build an AST for filter expressions and compile it on the server IIRC. The client gets filtered data from the server.




Guidelines | FAQ | Lists | API | Security | Legal | Apply to YC | Contact

Search: