zulip

Commit Graph

Author	SHA1	Message	Date
Alex Vandiver	5be7bc58fe	upload: Use content_disposition_header from Django 4.2. The code for this was merged in Django 4.2: https://code.djangoproject.com/ticket/34194	2023-05-11 14:51:28 -07:00
Alex Vandiver	04e7621668	upload: Rename upload_message_image_from_request. The table is named Attachment, and not all of them are images.	2023-03-02 16:36:19 -08:00
Alex Vandiver	23894fc9a3	uploads: Set Content-Type and -Disposition from Django for local files. Similar to the previous commit, Django was responsible for setting the Content-Disposition based on the filename, whereas the Content-Type was set by nginx based on the filename. This difference is not exploitable, as even if they somehow disagreed with Django's expected Content-Type, nginx will only ever respond with Content-Types found in `uploads.types` -- none of which are unsafe for user-supplied content. However, for consistency, have Django provide both Content-Type and Content-Disposition headers.	2023-02-07 17:12:02 +00:00
Alex Vandiver	2f6c5a883e	CVE-2023-22735: Provide the Content-Disposition header from S3. The Content-Type of user-provided uploads was provided by the browser at initial upload time, and stored in S3; however, `04cf68b45e` switched to determining the Content-Disposition merely from the filename. This makes uploads vulnerable to a stored XSS, wherein a file uploaded with a content-type of `text/html` and an extension of `.png` would be served to browsers as `Content-Disposition: inline`, which is unsafe. The `Content-Security-Policy` headers in the previous commit mitigate this, but only for browsers which support them. Revert parts of `04cf68b45e`, specifically by allowing S3 to provide the Content-Disposition header, and using the `ResponseContentDisposition` argument when necessary to override it to `attachment`. Because we expect S3 responses to vary based on this argument, we include it in the cache key; since the query parameter has dashes in it, we can't use use the helper `$arg_` variables, and must parse it from the query parameters manually. Adding the disposition may decrease the cache hit rate somewhat, but downloads are infrequent enough that it is unlikely to have a noticeable effect. We take care to not adjust the cache key for requests which do not specify the disposition.	2023-02-07 17:09:52 +00:00
Alex Vandiver	d41a00b83b	uploads: Extra-escape internal S3 paths. In nginx, `location` blocks operate on the _decoded_ URI[^1]: > The matching is performed against a normalized URI, after decoding > the text encoded in the “%XX” form This means that if a user-uploaded file contains characters that are not URI-safe, the browser encodes them in UTF-8 and then URI-encodes them -- and nginx decodes them and reassembles the original character before running the `location ~ ^/...` match. This means that the `$2` _is not URI-encoded_ and _may contain non-ASCII characters. When `proxy_pass` is passed a value containing one or more variables, it does no encoding on that expanded value, assuming that the bytes are exactly as they should be passed to the upstream. This means that directly calling `proxy_pass https://$1/$2` would result in sending high-bit characters to the S3 upstream, which would rightly balk. However, a longstanding bug in nginx's `set` directive[^2] means that the following line: ```nginx set $download_url https://$1/$2; ``` ...results in nginx accidentally URI-encoding $1 and $2 when they are inserted, resulting in a `$download_url` which is suitable to pass to `proxy_pass`. This bug is only present with numeric capture variables, not named captures; this is particularly relevant because numeric captures are easily overridden by additional regexes elsewhere, as subsequent commits will add. Fixing this is complicated; nginx does not supply any way to escape values[^3], besides a third-party module[^4] which is an undue complication to begin using. The only variable which nginx exposes which is _not_ un-escaped already is `$request_uri`, which contains the very original URL sent by the browser -- and thus can't respect any work done in Django to generate the `X-Accel-Redirect` (e.g., for `/user_uploads/temporary/` URLs). We also cannot pass these URLs to nginx via query-parameters, since `$arg_foo` values are not URI-decoded by nginx, there is no function to do so[^3], and the values must be URI-encoded because they themselves are URLs with query parameters. Extra-URI-encode the path that we pass to the `X-Accel-Redirect` location, for S3 redirects. We rely on the `location` block un-escaping that layer, leaving `$s3_hostname` and `$s3_path` as they were intended in Django. This works around the nginx bug, with no behaviour change. [^1]: http://nginx.org/en/docs/http/ngx_http_core_module.html#location [^2]: https://trac.nginx.org/nginx/ticket/348 [^3]: https://trac.nginx.org/nginx/ticket/52 [^4]: https://github.com/openresty/set-misc-nginx-module#set_escape_uri	2023-02-07 17:09:52 +00:00
Anders Kaseorg	81a7c7502f	requirements: Upgrade Python requirements. Signed-off-by: Anders Kaseorg <anders@zulip.com>	2023-02-03 16:36:54 -08:00
Anders Kaseorg	d3164016f5	ruff: Fix UP032 Use f-string instead of `format` call. Signed-off-by: Anders Kaseorg <anders@zulip.com>	2023-01-23 11:18:36 -08:00
Alex Vandiver	04cf68b45e	uploads: Serve S3 uploads directly from nginx. When file uploads are stored in S3, this means that Zulip serves as a 302 to S3. Because browsers do not cache redirects, this means that no image contents can be cached -- and upon every page load or reload, every recently-posted image must be re-fetched. This incurs extra load on the Zulip server, as well as potentially excessive bandwidth usage from S3, and on the client's connection. Switch to fetching the content from S3 in nginx, and serving the content from nginx. These have `Cache-control: private, immutable` headers set on the response, allowing browsers to cache them locally. Because nginx fetching from S3 can be slow, and requests for uploads will generally be bunched around when a message containing them are first posted, we instruct nginx to cache the contents locally. This is safe because uploaded file contents are immutable; access control is still mediated by Django. The nginx cache key is the URL without query parameters, as those parameters include a time-limited signed authentication parameter which lets nginx fetch the non-public file. This adds a number of nginx-level configuration parameters to control the caching which nginx performs, including the amount of in-memory index for he cache, the maximum storage of the cache on disk, and how long data is retained in the cache. The currently-chosen figures are reasonable for small to medium deployments. The most notable effect of this change is in allowing browsers to cache uploaded image content; however, while there will be many fewer requests, it also has an improvement on request latency. The following tests were done with a non-AWS client in SFO, a server and S3 storage in us-east-1, and with 100 requests after 10 requests of warm-up (to fill the nginx cache). The mean and standard deviation are shown. \| \| Redirect to S3 \| Caching proxy, hot \| Caching proxy, cold \| \| ----------------- \| ------------------- \| ------------------- \| ------------------- \| \| Time in Django \| 263.0 ms ± 28.3 ms \| 258.0 ms ± 12.3 ms \| 258.0 ms ± 12.3 ms \| \| Small file (842b) \| 586.1 ms ± 21.1 ms \| 266.1 ms ± 67.4 ms \| 288.6 ms ± 17.7 ms \| \| Large file (660k) \| 959.6 ms ± 137.9 ms \| 609.5 ms ± 13.0 ms \| 648.1 ms ± 43.2 ms \| The hot-cache performance is faster for both large and small files, since it saves the client the time having to make a second request to a separate host. This performance improvement remains at least 100ms even if the client is on the same coast as the server. Cold nginx caches are only slightly slower than hot caches, because VPC access to S3 endpoints is extremely fast (assuming it is in the same region as the host), and nginx can pool connections to S3 and reuse them. However, all of the 648ms taken to serve a cold-cache large file is occupied in nginx, as opposed to the only 263ms which was spent in nginx when using redirects to S3. This means that to overall spend less time responding to uploaded-file requests in nginx, clients will need to find files in their local cache, and skip making an uploaded-file request, at least 60% of the time. Modeling shows a reduction in the number of client requests by about 70% - 80%. The `Content-Disposition` header logic can now also be entirely shared with the local-file codepath, as can the `url_only` path used by mobile clients. While we could provide the direct-to-S3 temporary signed URL to mobile clients, we choose to provide the served-from-Zulip signed URL, to better control caching headers on it, and greater consistency. In doing so, we adjust the salt used for the URL; since these URLs are only valid for 60s, the effect of this salt change is minimal.	2023-01-09 18:23:58 -05:00
Alex Vandiver	58dc1059f3	uploads: Move unauth-signed tokens into view.	2023-01-09 18:23:58 -05:00
Alex Vandiver	ed6d62a9e7	avatars: Serve /user_avatars/ through Django, which offloads to nginx. Moving `/user_avatars/` to being served partially through Django removes the need for the `no_serve_uploads` nginx reconfiguring when switching between S3 and local backends. This is important because a subsequent commit will move S3 attachments to being served through nginx, which would make `no_serve_uploads` entirely nonsensical of a name. Serve the files through Django, with an offload for the actual image response to an internal nginx route. In development, serve the files directly in Django. We do _not_ mark the contents as immutable for caching purposes, since the path for avatar images is hashed only by their user-id and a salt, and as such are reused when a user's avatar is updated.	2023-01-09 18:23:58 -05:00
Alex Vandiver	f0f4aa66e0	uploads: Inline the one callsite of get_local_file_path. This helps make more explicit the assert_is_local_storage_path which makes using local_path safe.	2023-01-09 18:23:58 -05:00
Alex Vandiver	24f95a3788	uploads: Move internal upload serving path to under /internal/.	2023-01-09 18:23:58 -05:00
Alex Vandiver	cc9b028312	uploads: Set X-Accel-Redirect manually, without using django-sendfile2. The `django-sendfile2` module unfortunately only supports a single `SENDFILE` root path -- an invariant which subsequent commits need to break. Especially as Zulip only runs with a single webserver, and thus sendfile backend, the functionality is simple to inline. It is worth noting that the following headers from the initial Django response are _preserved_, if present, and sent unmodified to the client; all other headers are overridden by those supplied by the internal redirect[^1]: - Content-Type - Content-Disposition - Accept-Ranges - Set-Cookie - Cache-Control - Expires As such, we explicitly unset the Content-type header to allow nginx to set it from the static file, but set Content-Disposition and Cache-Control as we want them to be. [^1]: https://www.nginx.com/resources/wiki/start/topics/examples/xsendfile/	2023-01-09 18:23:58 -05:00
Alex Vandiver	679fb76acf	uploads: Provide our own Content-Disposition header. sendfile already applied a Content-Disposition header, but the algorithm may provide both `filename=` and `filename*=` values (which is potentially confusing to clients) and incorrectly slash-escapes quotes in Unicode strings. Django provides a correct implementation, but it is only accessible to FileResponse objects. Since the entire point is to offload the filehandle handling, we cannot use a FileResponse. Django 4.2 will make the function available outside of FileResponse. Until then, extract our own Content-Disposition handling, based on Django's. We remove the very verbose comment added in `d4360e2287`, describing Content-Disposition headers, as it does not add much.	2023-01-09 18:23:58 -05:00
Alex Vandiver	7c0d414aff	uploads: Split out S3 and local file backends into separate files. The uploads file is large, and conceptually the S3 and local-file backends are separable.	2023-01-09 18:23:58 -05:00
Lauryn Menard	aa796af0a8	upload: Remove `mimetype` url parameter in `get_file_info`. This `mimetype` parameter was introduced in `c4fa29a` and its last usage removed in `5bab2a3`. This parameter was undocumented in the OpenAPI endpoint documentation for `/user_uploads`, therefore there shouldn't be client implementations that rely on it's presence. Removes the `request.GET` call for the `mimetype` parameter and replaces it by getting the `content_type` value from the file, which is an instance of Django's `UploadedFile` class and stores that file metadata as a property. If that returns `None` or an empty string, then we try to guess the `content_type` from the filename, which is the same as the previous behaviour when `mimetype` was `None` (which we assume has been true since it's usage was removed; see above). If unable to guess the `content_type` from the filename, we now fallback to "application/octet-stream", instead of an empty string or `None` value. Also, removes the specific test written for having `mimetype` as a url parameter in the request, and replaces it with a test that covers when we try to guess `content_type` from the filename.	2022-08-08 16:06:09 -07:00
Zixuan James Li	f42465319b	upload: Refactor file size out of get_file_info. We have already checked the size of the file in `upload_file_backend`. This is the only caller of `upload_message_image_from_request`, and indirectly the only caller of `get_file_info`. There is no need to retrieve this information again. Signed-off-by: Zixuan James Li <p359101898@gmail.com>	2022-07-29 14:09:12 -07:00
Zixuan James Li	0ec561ab57	upload: Add assertions before accessing uploaded files. Signed-off-by: Zixuan James Li <p359101898@gmail.com>	2022-06-23 22:09:05 -07:00
Zixuan James Li	4cf3ba5744	typing: Fix typical typing typos. Signed-off-by: Zixuan James Li <p359101898@gmail.com>	2022-06-23 19:25:48 -07:00
Aman Agrawal	b799ec32b0	upload: Allow rate limited access to spectators for uploaded files. We allow spectators access to uploaded files in web public streams but rate limit the daily requests to 1000 per file by default.	2022-03-24 10:50:00 -07:00
Alex Vandiver	abed174b12	uploads: Add an endpoint which forces a download. This is most useful for images hosted in S3, which are otherwise always displayed in the browser.	2022-03-22 15:05:02 -07:00
Lauryn Menard	3be622ffa7	backend: Add request as parameter to json_success. Adds request as a parameter to json_success as a refactor towards making `ignored_parameters_unsupported` functionality available for all API endpoints. Also, removes any data parameters that are an empty dict or a dict with the generic success response values.	2022-02-04 15:16:56 -08:00
PIG208	dcbb2a78ca	python: Migrate most json_error => JsonableError. JsonableError has two major benefits over json_error: * It can be raised from anywhere in the codebase, rather than being a return value, which is much more convenient for refactoring, as one doesn't potentially need to change error handling style when extracting a bit of view code to a function. * It is guaranteed to contain the `code` property, which is helpful for API consistency. Various stragglers are not updated because JsonableError requires subclassing in order to specify custom data or HTTP status codes.	2021-06-30 16:22:38 -07:00
Anders Kaseorg	e7ed907cf6	python: Convert deprecated Django ugettext alias to gettext. django.utils.translation.ugettext is a deprecated alias of django.utils.translation.gettext as of Django 3.0, and will be removed in Django 4.0. Signed-off-by: Anders Kaseorg <anders@zulip.com>	2021-04-15 18:01:34 -07:00
Anders Kaseorg	6e4c3e41dc	python: Normalize quotes with Black. Signed-off-by: Anders Kaseorg <anders@zulip.com>	2021-02-12 13:11:19 -08:00
Anders Kaseorg	11741543da	python: Reformat with Black, except quotes. Signed-off-by: Anders Kaseorg <anders@zulip.com>	2021-02-12 13:11:19 -08:00
Anders Kaseorg	f364d06fb5	python: Convert percent formatting to .format for translated strings. Signed-off-by: Anders Kaseorg <anders@zulip.com>	2020-06-15 16:24:46 -07:00
Anders Kaseorg	365fe0b3d5	python: Sort imports with isort. Fixes #2665. Regenerated by tabbott with `lint --fix` after a rebase and change in parameters. Note from tabbott: In a few cases, this converts technical debt in the form of unsorted imports into different technical debt in the form of our largest files having very long, ugly import sequences at the start. I expect this change will increase pressure for us to split those files, which isn't a bad thing. Signed-off-by: Anders Kaseorg <anders@zulip.com>	2020-06-11 16:45:32 -07:00
Anders Kaseorg	67e7a3631d	python: Convert percent formatting to Python 3.6 f-strings. Generated by pyupgrade --py36-plus. Signed-off-by: Anders Kaseorg <anders@zulip.com>	2020-06-10 15:02:09 -07:00
Mateusz Mandera	4018dcb8e7	upload: Include filename at the end of temporary access URLs.	2020-04-20 10:25:48 -07:00
Tim Abbott	0ccc0f02ce	upload: Support requesting a temporary unauthenticated URL. This is be useful for the mobile and desktop apps to hand an uploaded file off to the system browser so that it can render PDFs (Etc.). The S3 backend implementation is simple; for the local upload backend, we use Django's signing feature to simulate the same sort of 60-second lifetime token. Co-Author-By: Mateusz Mandera <mateusz.mandera@protonmail.com>	2020-04-17 09:08:10 -07:00
rht	41e3db81be	dependencies: Upgrade to Django 2.2.10. Django 2.2.x is the next LTS release after Django 1.11.x; I expect we'll be on it for a while, as Django 3.x won't have an LTS release series out for a while. Because of upstream API changes in Django, this commit includes several changes beyond requirements and: * urls: django.urls.resolvers.RegexURLPattern has been replaced by django.urls.resolvers.URLPattern; affects OpenAPI code and related features which re-parse Django's internals. https://code.djangoproject.com/ticket/28593 * test_runner: Change number to suffix. Django changed the name in this ticket: https://code.djangoproject.com/ticket/28578 * Delete now-unnecessary SameSite cookie code (it's now the default). * forms: urlsafe_base64_encode returns string in Django 2.2. https://docs.djangoproject.com/en/2.2/ref/utils/#django.utils.http.urlsafe_base64_encode * upload: Django's File.size property replaces _get_size(). https://docs.djangoproject.com/en/2.2/_modules/django/core/files/base/ * process_queue: Migrate to new autoreload API. * test_messages: Add an extra query caused by .refresh_from_db() losing the .select_related() on the Realm object. * session: Sync SessionHostDomainMiddleware with Django 2.2. There's a lot more we can do to take advantage of the new release; this is tracked in #11341. Many changes by Tim Abbott, Umair Waheed, and Mateusz Mandera squashed are squashed into this commit. Fixes #10835.	2020-02-13 16:27:26 -08:00
Anders Kaseorg	4d49a20430	requirements: Upgrade django-sendfile2 from 0.4.3 to 0.5.1. The module was renamed from sendfile to django_sendfile. Signed-off-by: Anders Kaseorg <anders@zulipchat.com>	2020-02-05 12:38:10 -08:00
Tim Abbott	c869a3bf82	upload: Fix browser caching of uploads with local uploads backend. Apparently, our change in `b8a1050fc4` to stop caching responses on API endpoints accidentally ended up affecting uploaded files as well. Fix this by explicitly setting a Cache-Control header in our Sendfile responses, as well as changing our outer API caching code to only set the never cache headers if the view function didn't explicitly specify them itself. This is not directly related to #13088, as that is a similar issue with the S3 backend. Thanks to Gert Burger for the report.	2019-10-01 15:15:17 -07:00
Anders Kaseorg	780ecb672b	CVE-2019-16216: Fix MIME type validation. * Whitelist a small number of image/ types to be served as non-attachments. * Serve the file using the type that we validated rather than relying on an independent guess to match. This issue can lead to a stored XSS security vulnerability for older browsers that don't support Content-Security-Policy. It primarily affects servers using Zulip's local file uploads backend for servers running Ubuntu 16.04 Xenial or newer; the legacy local file upload backend for (now EOL) Ubuntu 14.04 Trusty was not affected and it has limited impact for the S3 upload backend (which uses an unprivileged S3 bucket domain to serve files). Signed-off-by: Anders Kaseorg <anders@zulipchat.com>	2019-09-11 15:46:36 -07:00
Anders Kaseorg	72655611ce	requirements: Use maintained fork django-sendfile2 of django-sendfile The original seems to be unmaintained (johnsensible/django-sendfile#65). Notably, this fixes a bug in the filename parameter, which perviously showed the Python 3 repr of a byte string (johnsensible/django-sendfile#49). Signed-off-by: Anders Kaseorg <anders@zulipchat.com>	2019-08-12 15:40:08 -07:00
Anders Kaseorg	61982d9d47	uploads: Revert "Url encoded name of the file should be an ascii." This reverts commit `fd9dd51d16` (#1815). The issue described does not exist in Python 3, where urllib.parse now _only_ accepts (Unicode) str and does the right thing with it. The workaround was not being triggered and would have failed if it were. Signed-off-by: Anders Kaseorg <anders@zulipchat.com>	2019-04-22 22:28:39 -07:00
Anders Kaseorg	4e21cc0152	views: Remove unused imports. Signed-off-by: Anders Kaseorg <andersk@mit.edu>	2019-02-02 17:23:43 -08:00
Tim Abbott	5f7691b74e	upload: Remove unnecessary use of has_request_variables. All the parameters for this function are parsed in urls.py.	2018-07-01 01:47:03 -07:00
Aditya Bansal	d4360e2287	uploads: Make django-sendfile to force downloading attachments. We start to force downloads for the attachment files. We do this for all files except images or pdf's. We would like images or pdf's to open up in browser itself. Tweaked by tabbott for comment clarity and correctness.	2018-03-14 11:22:10 -07:00
Aditya Bansal	efe8545303	local-uploads: Start running authentication checks on file requests. From here on we start to authenticate uploaded file request before serving this files in production. This involves allowing NGINX to pass on these file requests to Django for authentication and then serve these files by making use on internal redirect requests having x-accel-redirect field. The redirection on requests and loading of x-accel-redirect param is handled by django-sendfile. NOTE: This commit starts to authenticate these requests for Zulip servers running platforms either Ubuntu Xenial (16.04) or above. Fixes: #320 and #291 partially.	2018-02-16 05:06:37 +05:30
Vishnu Ks	43a6439b3b	upload: Enforce per-realm quota.	2018-01-29 16:06:11 -08:00
Greg Price	55cf54c087	upload: Remove old per-user quota feature. We'll replace this primarily with per-realm quotas (plus the simple per-file limit of settings.MAX_FILE_UPLOAD_SIZE, 25 MiB by default). We do want per-user quotas too, but they'll need some more management apparatus around them so an admin has a practical way to set them differently for different users. And the error handling in this existing code is rather confused. Just clear this feature out entirely for now; then we'll build the per-realm version more cleanly, and then we can later add back per-realm quotas modelled after that. The migration to actually remove the field is in a subsequent commit. Based in part on work by Vishnu Ks (hackerkid).	2018-01-29 16:06:11 -08:00
rht	e538f4dd44	zerver/views: Use Python 3 syntax for typing. Edited by tabbott to remove state.py and streams.py, because of problems with the original PR's changes, and wrap some long lines.	2017-11-27 17:10:39 -08:00
rht	88a828dd0c	Remove six.moves.urllib usage.	2017-11-09 10:00:00 -08:00
Greg Price	68b0a419ec	decorator: Cut a bunch of dead imports of two view decorators. Saw these when grepping for these two decorators; they're actually more numerous than the surviving use sites are. Cut out the noise.	2017-11-04 19:27:00 -07:00
rht	15ca13c8de	zerver/views: Remove absolute_import.	2017-09-27 10:00:39 -07:00
Tim Abbott	6a50e13156	uploads: Remove legacy /json/upload_file endpoint. This migrates Zulip to use the equivalent API endpoint that has been present for a while.	2017-07-31 13:08:06 -07:00
Ethan	d4d689532d	mypy: serve_local return type to FileResponse.	2017-05-25 15:41:52 -07:00
rahuldeve	60803137f2	uploads: Add authorization check before serving files. This is a remerge of `e985b57259` (after resolving merge conflicts, updating the tests, adding mypy annotations etc.), which should now be correct, because we've done the necessary database migration. The rebase/remerge work was done by Tim Abbott and Aditya Bansal. This is an important part of #320.	2017-04-07 16:35:28 -07:00

1 2

70 Commits