THE EXTENSION
It reads the tab you're already in.
Select a passage or take the whole page. The extension extracts and hands off — it never becomes a second app you have to remember to open, and it never carries the heavy work.
The two most commonly used attention functions are additive attention and dot-product attention. Dot-product attention is identical to our algorithm, except for the scaling factor. We suspect that for large values of d, the dot products grow large in magnitude, pushing the softmax function into regions where it has extremely small gradients. To counteract this effect, we scale the dot products.