feat: add node_readiness_rules with enforcement_mode and dry_run labels - #448
feat: add node_readiness_rules with enforcement_mode and dry_run labels#448rawadhossain wants to merge 1 commit into
Conversation
|
[APPROVALNOTIFIER] This PR is NOT APPROVED This pull-request has been approved by: rawadhossain The full list of commands accepted by this bot can be found here. DetailsNeeds approval from an approver in each of these files:Approvers can indicate their approval by writing |
✅ Deploy Preview for node-readiness-controller ready!
To edit notification comments on pull requests, go to your Netlify project configuration. |
|
Hi @rawadhossain. Thanks for your PR. I'm waiting for a kubernetes-sigs member to verify that this patch is reasonable to test. If it is, they should reply with Tip We noticed you've done this a few times! Consider joining the org to skip this step and gain Once the patch is verified, the new status will be reflected by the I understand the commands that are listed here. DetailsInstructions for interacting with me using PR comments are available here. If you have questions or suggestions related to my behavior, please file an issue against the kubernetes-sigs/prow repository. |
|
/cc @ajaysundark |
|
/cc @AvineshTripathi |
|
/ok-to-test |
|
/retest |
AvineshTripathi
left a comment
There was a problem hiding this comment.
Since we're already calling ListRules in the collector, we can compute the enforcement_mode/dry_run breakdown there instead of maintaining a separate in-memory sync. The local ruleCache can lag behind actual cluster state and can be unreliable.
cc @ajaysundark wdyt?
58c5646 to
3ac4f69
Compare
|
@AvineshTripathi Good point. I tested both ways and found @ajaysundark since this metric wasn't part of the collector scope, should we switch or keep it separate for now? |
Signed-off-by: Rawad Hossain <rawad.hossain00@gmail.com>
3ac4f69 to
81c256d
Compare
I would prefer testing this delay on the scale before doing the scale. Can we run the scan? |
|
Yes we can and thanks it helped. Ran the scale test. Used the ruleCache staleness (current one):
Single changes are fine, but bigger difference shows up when a large number of rules change together, delay grows quite long. Scrape-time approach:
Since the collector already does that fetch, adding this one on barely adds anything. Based on these, would it make more sense to go with collector instead, specially for scale? Staleness gets pretty noticeable with bulk changes, also extra cost in the collector is very small. @AvineshTripathi @ajaysundark wdyt? I can make the changes if we decide to go with collector approach. |
Description
node_readiness_rules{enforcement_mode, dry_run}while keepingnode_readiness_rules_totalunchanged for compatibility.monitoring.mdanddocs/TEST_README.md.from the design doc:
Related to #446
Type of Change
/kind feature
Testing
make test,make lint,go vet,gofmt -lall pass.Checklist
make testpassesmake lintpasses