zulip

Commit Graph

Author	SHA1	Message	Date
rht	b557b02f2f	analytics/lib: Remove unused imports (F401).	2017-11-07 16:37:07 -08:00
rht	ec5120e807	refactor: Remove six.moves.zip import.	2017-11-07 10:46:42 -08:00
rht	5cfffb0e51	analytics: Remove inheritance from object.	2017-11-06 08:53:48 -08:00
rht	dcc831f767	refactor: Replace all __unicode__ method with __str__. Close #6627.	2017-11-02 11:01:47 -07:00
rht	691598a88b	py3: Remove "from six.moves import range". This is no longer required, since in Python 3, this is what the range built-in does.	2017-10-17 23:28:14 -07:00
rht	2f3ae84e5a	py3: Remove all `__future__ import division`.	2017-10-17 23:09:12 -07:00
rht	a603a4f9f5	Remove `from __future__ import absolute_import`. Except in: - docs/writing-bots-guide.md, because bots are supposed to be Python 2 compatible - puppet/zulip_ops/files/zulip-ec2-configure-interfaces, because this script is still on python2.7 - tools/lint - tools/linter_lib - tools/lister.py For the latter two, because they might be yanked away to a separate repo for general use with other FLOSS projects.	2017-10-17 22:59:42 -07:00
Rishi Gupta	c7bdabbda8	analytics: Disallow non-UTC fill times in process_count_stat. No change in behavior, but we aren't supporting non-UTC times in analytics as a whole any more, so might as well change this check as well.	2017-10-05 11:22:06 -07:00
Rishi Gupta	0596c4a810	analytics: Enforce various datetime arguments are in UTC. Sort of a hacky hammer, but * The original design of the analytics system mistakenly attempted to play nicely with non-UTC datetimes. * Timezone errors are really hard to find and debug, and don't jump out that easily when reading code. I don't know of any outstanding errors, but putting a few "assert this timezone is in UTC" around will hopefully reduce the chance that there are any current or future timezone errors. Note that none of these functions are called outside of the analytics code (and tests). This commit also doesn't change any current behavior, assuming a database where all datetimes have been being stored in UTC.	2017-10-05 11:22:06 -07:00
Rishi Gupta	0f31cddf49	analytics: Add management command to clear single stat.	2017-10-05 11:22:06 -07:00
Aditya Bansal	d9c9bfe7f6	logger: Add new create_logger abstraction to simplify logging. This deduplicates a ton of Python logger-creation code to use a single standard implementation, so we can avoid copy-paste problems.	2017-08-27 18:31:53 -07:00
umkay	d9b23b39d3	mypy: Fix strict-optional in analytics.	2017-05-26 15:39:39 -07:00
Aditya Bansal	27b87943af	pep8: Add compliance with rule E261 to counts.py.	2017-05-07 23:21:50 -07:00
Rishi Gupta	61bf445da4	analytics: Restrict fill_to_time to hour boundaries in process_count_stat.	2017-04-28 16:15:07 -07:00
Rishi Gupta	5e49da9285	analytics: Only update daily stats on day boundaries. Previously we would update FillState for daily stats on hourly boundaries as well. This would create two extra queries on the FillState table every hour (for each CountStat), which adds roughly 50ms of extra processing for each CountStat each day, as well as two extra lines each hour in the analytics log. This can be a minor annoyance when backfilling stats.	2017-04-18 11:02:51 -07:00
Rishi Gupta	c5f1398052	analytics: Add section comments in counts.count_stats_. Also reorders the stats a bit.	2017-04-18 11:02:51 -07:00
Rishi Gupta	b335ad2794	models: Add MIN_INTERVAL_LENGTH to UserActivityInterval. Was previously a floating magic number appearing in both zerver/lib/actions.py and analytics/lib/counts.py.	2017-04-18 11:02:51 -07:00
hackerkid	5c8f011d66	Remove unused timezone import.	2017-04-16 12:28:56 -07:00
Rishi Gupta	49bd330304	analytics: Add class DependentCountStat and stat realm_active_humans::day.	2017-04-14 11:41:07 -07:00
Rishi Gupta	1e8d2b984d	counts.py: Rename DataCollector-level operations to be more generic. We're about to use these for DependentCountStats that will run SQL queries on the analytics tables instead of the zerver tables.	2017-04-14 11:41:07 -07:00
Rishi Gupta	47cf1d15ba	counts.py: Move performance logging call out of pull_functions. Makes it less likely someone will write a pull function in the future and forget.	2017-04-14 11:41:07 -07:00
Rishi Gupta	6dff22cbaf	counts.py: Change check for LoggingCountStat to use isinstance. I think this is more pythonic? We could also get rid of LoggingCountStats altogether, since it's now just a special case of CountStat (is_logging == data_collector.pull_function is None). But I think it's nice to keep the distinction since they behave so differently.	2017-04-14 11:41:07 -07:00
Rishi Gupta	b45185562a	counts.py: Fix out of date comments.	2017-04-14 11:41:07 -07:00
Rishi Gupta	ac2cc9e2da	counts.py: Reorganize file into logical sections. No changes to code or behavior.	2017-04-14 11:41:07 -07:00
Rishi Gupta	50868b98a9	counts.py: Change pull_function to take a property instead of a full stat. Removes the circular dependency of CountStat containing a DataCollector, and DataCollector containing a function that takes a CountStat as an argument.	2017-04-14 11:41:07 -07:00
Rishi Gupta	eadfc743c8	counts.py: Remove CustomPullCountStat.	2017-04-14 11:41:07 -07:00
Rishi Gupta	118b44d4f0	counts.py: Change DataCollector to take a pull_function argument. This will allow us to appropriately generalize CountStat to include LoggingCountStat and CustomPullCountStat. It'll also make life easier when we introduce DependentCountStat.	2017-04-14 11:41:07 -07:00
Rishi Gupta	f9e56ad25d	counts.py: Move DataCollector declarations into CountStat declarations. The previous zerver_* names were unwieldy and not very readable. This also puts more of the useful information in one place; in particular, makes it easier to skim a CountStat declaration and see if we're collecting it at a user/stream granularity or a realm granularity.	2017-04-14 11:41:07 -07:00
Rishi Gupta	c20e79ab1f	counts.py: Rename DataCollector.analytics_table to output_table.	2017-04-14 11:41:07 -07:00
Rishi Gupta	6369d23633	counts.py: Rename ZerverCountQuery to DataCollector. Not the final form of DataCollector, but the name change causes a big diff so separating it out.	2017-04-14 11:41:07 -07:00
Rishi Gupta	b3991e2557	counts.py: Move CountStat.group_by into ZerverCountQuery. Part of a larger refactoring to reduce cyclic dependencies between CountStat and DataCollector (coming soon).	2017-04-14 11:41:07 -07:00
Rishi Gupta	341e1b54fc	counts.py: Remove zerver_table from ZerverCountQuery. Was only needed for filter_args, which are now gone.	2017-04-14 11:41:07 -07:00
Rishi Gupta	661de6bf25	counts.py: Remove filter_args argument from CountStat definition. It turned out to not be that useful once we added subgroup. The previous design of the CountStat object also assumed more reuseability of the _query strings than what ended up happening. The filter_args also had some carrying costs: It's hard to be confident that filter_args other than the ones explicitly in our tests would have had expected behavior. * The filter_args/join_args system is the most complex part of the CountStat object, and makes understanding the *_query strings unnecessarily difficult for a new contributor.	2017-04-14 11:41:07 -07:00
Rishi Gupta	4dfadba244	counts.py: Hardcode is_active=true in count_user_by_realm_query. A step towards removing filter_args from the CountStat object.	2017-04-14 11:41:07 -07:00
Rishi Gupta	6bb97db136	analytics: Add active_users_audit:is_bot:day.	2017-04-14 11:41:07 -07:00
Rishi Gupta	cc75d83b74	counts.py: Reorder count_stats_ to put similar stats together.	2017-04-14 11:41:07 -07:00
Rishi Gupta	2f74ccabf9	analytics: Add 15day_actives CountStat.	2017-04-14 11:41:07 -07:00
Rishi Gupta	9b661ca91f	analytics: Replace CountStat.is_gauge with interval. Groundwork for allowing stats like "Monthly Active Users". CountStat.interval is no longer as clean a value as before, so removed it from views.get_chart_data. It wasn't being used by the frontend anyway. Removing interval from logger calls in counts.py is not a big loss since we now include the frequency (which is typically also the interval) in CountStat.property.	2017-04-14 11:41:07 -07:00
Rishi Gupta	d6c5c672d3	analytics: Add minutes_active CountStat.	2017-04-14 11:41:07 -07:00
hollywoodno	dd067c761a	analytics: Separate private messages from group private messages. This makes it possible for our graphs to show the group private message counts as separate from 1:1 private messages. Fixes #4102.	2017-03-20 11:46:29 -07:00
Rishi Gupta	7c6f0033ed	analytics: Add test for do_drop_all_analytics_tables.	2017-03-14 16:59:54 -07:00
Rishi Gupta	87981a2bf1	analytics: Fix direct import of models in migrations.	2017-03-14 16:59:54 -07:00
Rishi Gupta	ebebd04587	analytics: Fix ValueErrors affecting test coverage. Pathways that only catch internal code errors should use AssertionError so that they are not included when computing test coverage.	2017-03-14 16:59:54 -07:00
Rishi Gupta	b18bfe6771	analytics: Standardize format of zerver count queries. count_message_type_by_user_query is in a different format (no WHERE clause) from the rest since I'm having a hard time reasoning about how that would interact with the LEFT JOIN, especially given that there are %(join_args)s.	2017-03-14 16:59:54 -07:00
Rishi Gupta	8feea6c598	analytics: Add LoggingCountStat for number of users.	2017-03-04 16:46:09 -08:00
Raghav Jajodia	a3a03bd6a5	mypy: Added Dict, List and Set imports. Fixed mypy errors associated with the upgrade.	2017-03-04 14:33:44 -08:00
Rishi Gupta	8bea47d6b5	analytics: Do a stylistic cleanup of TestProcessCountStat.	2017-03-03 16:12:12 -08:00
Rishi Gupta	6c784d6321	analytics: Refactor COUNT_STATS declaration to not repeat itself.	2017-03-03 16:11:28 -08:00
Rishi Gupta	20255e48a4	analytics: Change messages_sent_to_stream to a daily stat. Analytics database tables are getting big, and so we're likely moving to a model where ~all stats are day stats, and we keep hourly stats only for the last N days. Also changed the name because: * messages_sent_* suggests the counts (summed over subgroup) should be the same as the other messages_sent stats, but they are different (these don't include PMs). * messages_sent_by_stream:is_bot:day is longer than 32 characters, the max allowable length for a BaseCount.property. Includes a database migration to remove the old stat from the analytics tables.	2017-03-03 16:11:28 -08:00
Rishi Gupta	5eb5fa3f31	analytics: Change time_range to not include current day/hour. Current day/hour will always be 0, since we haven't computed it yet for the CountStat tables.	2017-02-02 10:59:52 -08:00

1 2

95 Commits