1. Match patterns
A match pattern is a string used to match URLs.
A match pattern is either the special value <all_urls>, or a string of the
form:
<scheme>://<host><path>
A match pattern a subsumes a match pattern b if every URL matched by b is also matched by a.
1.1. Grammar
1.1.1. scheme
scheme is one of a fixed list of literal scheme names, or *.
A pattern’s scheme is the substring of the match pattern before the
first ://.
The literal scheme values http, https, file, and ftp, plus a browser’s own extension scheme, are valid scheme values in Chrome, Firefox, and Safari alike. http, https,
file, and ftp are four of the URL Standard’s six special schemes; the other two are ws
and wss. data is not a special scheme, which matters for
§ 1.2 <all_urls>: Firefox is the only one of the three browsers that includes data in
<all_urls>.
http, https, file, ftp, and a browser’s own extension scheme,
does a match pattern’s scheme component accept? Firefox additionally accepts ws, wss, and
data. Chrome additionally accepts ws/wss for host_permissions (but not for
content_scripts.matches) and, only in a non-default developer configuration, its own
chrome:// scheme. Safari accepts none of ws, wss, data, or an equivalent internal-pages
scheme in a match pattern at all.
If scheme is *, the pattern matches a fixed, implementation-defined
set of schemes in place of a literal scheme. That set always includes http and https in
Chrome, Firefox, and Safari alike.
Should the * scheme wildcard expand to include ws/wss, matching Firefox, or only
http/https, matching Chrome and Safari?
1.1.2. host
host is one of:
-
*, matching any host. -
*.followed by a domain suffix (containing no*or/), matching that domain and any of its subdomains. -
a literal domain or IP-address literal (containing no
*or/), matching only that exact host.
Those are the only shapes host takes when it contains a * at all; see
§ 1.3 Parsing a match pattern for what a * in any other position does to
parsing.
When scheme is file, host is omitted (the pattern has the form
file://<path> or file:///<path>, with <path> beginning at the third
slash regardless of what, if anything, appears between the second and
third slash).
Host comparison is case-insensitive.
Chrome and Safari match hosts case-insensitively. Firefox does not: it stores the pattern’s
host verbatim and compares it byte-for-byte, so *://EXAMPLE.com/* matches nothing there.
Specifying case-insensitive host matching is a behavior change for Firefox, not a documentation
fix.
Only Chrome does this today. Firefox and Safari store the host verbatim, so a pattern host written with literal Unicode matches nothing in those two browsers: the navigated URL’s host is always Punycode-encoded, and the author has to write the Punycode form themselves.
A port is never written in host; see § 1.3 Parsing a match pattern for what writing one does to parsing.
Does a match pattern need a port component at all? Chrome’s
grammar has one: a literal <host>:<port> matches only that port, and
omitting the port (the common case) matches any port. Safari’s and
Firefox’s grammars have none.
https://*.example.com/* matches https://example.com/,
https://www.example.com/, and https://a.b.example.com/, but not
https://notexample.com/.
1.1.3. path
path begins with /; see § 1.3 Parsing a match pattern for what a
pattern with no path does to parsing.
http://example.com/ has a path of / and is a valid match pattern.
http://example.com, with no trailing slash and therefore no path at
all, fails to parse.
path may contain any number of * characters, each matching zero or
more characters. No other character in path has special meaning; in
particular, unlike a glob, ? is an ordinary literal character in a
match pattern’s path, not a single-character wildcard.
Path matching is case-sensitive in Chrome, Firefox, and Safari (Chrome exposes a case-insensitive matching mode in its API, but no call site in the browser’s own permission or content-script matching code uses it).
Chrome and Firefox match it against the URL’s path and query string together (that is,
against pathname + search). Safari matches it against the path alone; the query
string is excluded.
Chrome decodes percent-encoded octets in both the URL and the pattern’s path before comparing
(falling back to a raw comparison for byte sequences that aren’t valid UTF-8). Firefox and
Safari perform no decoding at all: a pattern containing a literal, unescaped reserved character
will not match a URL where the user agent encoded that same character, in either browser. See
also w3c/webextensions#945, a related but
distinct percent-encoding discussion scoped to declarativeNetRequest’s urlFilter syntax
rather than match patterns.
https://example.com/foo/* matches https://example.com/foo/,
https://example.com/foo/bar, and https://example.com/foo/bar?baz,
but not https://example.com/foobar or https://example.com/foo
(without a trailing slash).
1.2. <all_urls>
The special match pattern <all_urls> matches all URLs whose scheme is
in a permitted set of schemes, regardless of host or path. http and https are in that set
in every engine, in every context that defines one.
Beyond http and https, each browser permits a different subset of the URL Standard’s six
special schemes plus data for <all_urls>:
| scheme | Chrome | Firefox | Safari |
|---|---|---|---|
http
| yes | yes | yes |
https
| yes | yes | yes |
file
| yes | yes | opt-in |
ftp
| yes | yes | no |
ws, wss
| host_permissions only
| yes | no |
data
| no | yes | no |
Safari’s <all_urls> also includes its own extension scheme, shared by all of its
extensions. Chrome’s and Firefox’s do not.
<all_urls> cover a single portable scheme set across browsers, or leave file: and
other non-http(s) schemes as a browser-specific opt-in (Safari’s model) rather than an
unconditional inclusion (Firefox’s model)?
content_scripts.matches list of ["<all_urls>"]:
| URL | Chrome | Firefox | Safari |
|---|---|---|---|
https://example.com/
| runs | runs | runs |
ftp://example.com/
| runs | runs | does not run |
data:text/html,...
| does not run | runs | does not run |
file:///home/user/x.html
| runs only if file access has separately been granted | runs only if file access has separately been granted | runs only if file access has separately been granted |
1.3. Parsing a match pattern
To parse a match pattern given input:
-
If input is
<all_urls>, return<all_urls>. -
If input does not contain
://, return failure. -
Let scheme be the substring of input before the first
://. -
If scheme is not
*, and is not one of the permitted literal scheme values (see § 1.1.1 scheme), return failure. -
Let rest be the substring of input following that
://. -
If scheme is
file:-
If rest contains no U+002F (/), return failure.
-
Let host be the empty string. Any text of rest before its first U+002F (/) is discarded, not read as a host.
-
Let path be the substring of rest starting at its first U+002F (/), inclusive.
-
-
Otherwise:
-
If rest does not contain U+002F (/), return failure.
-
Let host be the substring of rest before its first U+002F (/).
-
If host contains U+003A (:), return failure.
-
If host is not
*, host does not consist of*.followed by one or more characters containing no U+002A (*), and host contains U+002A (*), return failure. -
Let path be the substring of rest starting at its first U+002F (/), inclusive.
-
-
Return a match pattern whose scheme is scheme, whose host is host, and whose path is path.
A file match pattern’s host is never read as a host at all: whatever
text appears between file:// and the next /, if any, is discarded
rather than validated. § 1.4 Matching algorithm separately
skips path comparison for a file pattern.
<all_urls> and https://*.example.com/* both parse successfully.
example.com/* (no ://) and https://example.com (no path) both
fail to parse. http://example.com:8080/* also fails to parse; see
§ 1.1.2 host for whether match patterns take a port
component.
1.4. Matching algorithm
pattern here is either <all_urls> or the result of a successful
parse of a match pattern string; see
§ 1.3 Parsing a match pattern. surface identifies which manifest key or API call is making the
check (for example content_scripts.matches or host_permissions); see
§ 1.1.1 scheme and § 1.2 <all_urls> for how the permitted and *-expansion scheme sets
vary by surface.
To determine whether a match pattern pattern matches a URL url for a manifest surface surface:
-
Let url record be the result of parsing url.
-
If pattern is
<all_urls>:-
If the scheme of url record is in the permitted scheme set for
<all_urls>given surface (see § 1.2 <all_urls>), return true. -
Otherwise, return false.
-
-
If pattern’s scheme is
*:-
If the scheme of url record is not one of the schemes
*expands to for surface, return false.
-
-
Otherwise:
-
If the host of url record does not match pattern’s host, return false.
-
If pattern’s scheme is not
file, and pattern’s path does not match the path of url record (see § 1.1.3 path for what "path" includes), return false. -
Return true.
This algorithm is run against the URL produced by
Determine the URL for matching a document when matching a document
for content_scripts, not necessarily against the document’s literal
URL; it does not itself define how the URL of a document with an
opaque origin is resolved.
Firefox has a stricter mode for host permissions in which a pattern with a wildcard host
(*.example.com, or a bare *) never matches. Chrome and Safari have no equivalent.
*://*.example.com/*:
| URL | Chrome | Firefox | Safari |
|---|---|---|---|
https://www.example.com/
| matches | matches | matches |
http://example.com/
| matches | matches | matches |
wss://example.com/
| does not match | matches | does not match |
Firefox’s * scheme wildcard additionally covers ws/wss.
1.5. Unparseable match patterns
A match pattern string that does not parse does not match any URL.
For a permission-bearing key (permissions, host_permissions,
optional_permissions, optional_host_permissions), an unparseable
entry is dropped and the extension loads with the rest of the list.