How do filters work? They seem pretty difficult implementation-wise since you can write them in any of the language bindings. My first guess is that you pipe all the data in a table to the client, and the client itself does the filtration. But this would be extraordinarily inefficient.
Piping all the data to the client would be extremely inefficient. Fortunately we don't do that.
When a filter is written in the client language it gets compiled into a protocol buffer which is sent to the cluster. This gets compiled into a query which is sent to each of the relevant shards for the table. This query has the filter baked right into it. The shards then go through their local copy of the data and filter out the rows which do not meet the query predicate. This data gets returned to the coordinating node and eventually to the user. Thus only the data the will actually be returned is ever transferred over the network.
Furthermore this process is done lazily. On the client side rather than getting back a huge array with the results of your filter you get back an iterator. This iterator stores a buffer of data which will be refilled as it is incremented.
To add to jdoliner's answer, the reason why you can write table('foo').filter(lambda x: x['bar'] > 5).run() is because we do some language trickery on the client side to compile the query to an AST. In this case, we overload greater than operator, call the lambda function once on the client with a special object, and return an AST. This AST is then sent to the server and executed there.
It's rather difficult to integrate into a host language like that smoothly from the driver implementation perspective, but once the driver is written the user experience is amazing because you can write queries that look exactly like Python, but they're executed entirely on the server.
Lambda is necessary because when you do nested subqueries, saying r['x'] is ambiguous and can cause all sorts of unpleasantness. So, if you use nested queries, the server rejects the implicit syntax and requires the use of lambda.
Lambda syntax is really nice too, I actually prefer it for writing queries.
I find the lambda trick not explicit and obvious enough. I fear I would do something stupid like trigger a side-effect without realising.
Again taking an example from SQLAlchemy, you can explicitly make a subquery, and then reference it instead of the original Table. A binding more like SQLAlchemy can probably written for RethinkDB.
No. E.g., you could write r.table('foo').update(lambda row: row.merge({'bar': row['bar'] + 1 })). A shortcut for this is r.table('foo').update({'bar': r['bar'] + 1 }). Neither is referentially transparent, and both work.
I believe both of those functions are referential transparent because they're pure functions. An example of a non referential transparent function would be: