node-feature-discovery

mirror of https://github.com/kubernetes-sigs/node-feature-discovery.git synced 2025-03-05 16:27:05 +00:00

Author	SHA1	Message	Date
Markus Lehtonen	ba4cebb29e	nfd-master: stop node-updater pool before reconfiguring api-controller Prevents potential race between node-updater pool and the api-controller when re-configuring nfd-master. Reconfiguration causes a new api-controller instance to be created so nfd api lister might change in the midst of processing a node update (if the pool was running). No actual issues related to this have been identified but races (like this) should still be avoided.	2024-04-09 10:45:07 +03:00
Markus Lehtonen	8709cccf71	nfd-master: parse kubeconfig even with NoPublish set Don't try to be too smart when kubeconfig is needed. In practice, the nfd-master really doesn't work anymore (with the NodeFeature API enabled) without a kubeconfig set. This patch fixes crashes happening when NoPublish is enabled, e.g. in listing all nodes in the nfd api handler and in getting single node objects in the node updater pool. This patch changes the kubeconfig parsing to happen at the creation of the nfd-master instance. We don't need to do that at reconfigure time as none of the dynamic config options affect it. Unit tests are adjusted, accordingly.	2024-04-08 14:25:27 +03:00
Markus Lehtonen	fcb8d3cda4	nfd-master: implement opts for modifying NfdMaster instance This provides a more controlled way for setting up the NfdMaster instance for testing.	2024-04-05 20:21:19 +03:00
Kubernetes Prow Robot	199d665046	Merge pull request #1656 from marquiz/devel/channel-simplify Tidy up usage of channels for signaling	2024-04-05 07:51:34 -07:00
Kubernetes Prow Robot	cb24f7c234	Merge pull request #1657 from marquiz/devel/master-label-whitelist nfd-master: prevent crash on empty config struct	2024-04-05 05:36:52 -07:00
Markus Lehtonen	26a80cf142	Tidy up usage of channels for signaling This started as a small effort to simplify the usage of "ready" channel in nfd-master. It extended into a wider simplification/unification of the channel usage.	2024-04-05 14:39:58 +03:00
Markus Lehtonen	b27676451a	nfd-master: prevent crash on empty config struct Change the handling of LabelWhiteList config option to use a pointer to detect when the option is unset. This doesn't fix any detected crash but is merely general improvement and stabilization, serving easier testing. Also, use the regexp type from the core libs for the config struct - dropping the unmasrhalling code for our custom regexp type - as the core regexp now implements unmarshaller itself.	2024-04-05 14:19:44 +03:00
Kubernetes Prow Robot	ad96c301a4	Merge pull request #1642 from marquiz/devel/master-updater-pool-lock nfd-master: protect node updater pool queueing with a lock	2024-04-05 03:31:10 -07:00
Markus Lehtonen	44a5a5b4a8	nfd-master: get node object only once when updating node Prevent excess queries of node objects from the Kubernetes apiserver. This significantly speeds up node updates (and reduces the load on the apiserver) as the client-side throttling (which is good) does not bite us that hard.	2024-04-04 14:44:52 +03:00
Oleg Zhurakivskyy	f2e9557a2d	nfd-topology-updater: Add liveness probe Signed-off-by: Oleg Zhurakivskyy <oleg.zhurakivskyy@intel.com>	2024-04-03 13:15:54 +03:00
Markus Lehtonen	bce446c5b6	nfd-master: protect node updater pool queueing with a lock Prevents races when (re-)starting the queue. There are no reports on issues related to this (and I haven't come up with any actual failure path in the current code) but better to be safe and follow the best practices.	2024-03-27 16:53:34 +02:00
Markus Lehtonen	c4e010eafd	nfd-master: do nfd API scheme registration in an init function Prevents (rare) races on nfd-master reconfigurartion. Previously the scheme was registered at nfd API controller creation/startup time. This caused a race with some lister/informer goroutines of the previous (stoppped) controller still running and accessing (reading) the sceme while we were updating (writing) it.	2024-03-27 15:26:16 +02:00
Oleg Zhurakivskyy	7bd27c757a	topology-updater: Set APIVersion, Kind in the OwnerReference explicitly APIVersion and Kind are empty in the returned namespace object and need to be set explicitly. Signed-off-by: Oleg Zhurakivskyy <oleg.zhurakivskyy@intel.com>	2024-03-20 20:09:06 +02:00
Kubernetes Prow Robot	0ad5e50f24	Merge pull request #1609 from ozhuraki/worker-health nfd-worker: Add liveness probe	2024-03-19 06:57:23 -07:00
Oleg Zhurakivskyy	8b63d17af7	nfd-worker: Add liveness probe Signed-off-by: Oleg Zhurakivskyy <oleg.zhurakivskyy@intel.com>	2024-03-19 15:34:53 +02:00
Kubernetes Prow Robot	c4ff25de52	Merge pull request #1596 from marquiz/devel/master-infinite-retry nfd-master: retry node updates indefinitely	2024-03-19 04:00:50 -07:00
Kubernetes Prow Robot	7df0f17f68	Merge pull request #1602 from ozhuraki/nrt-owner-ref Add owner reference to NRT object	2024-03-19 01:12:59 -07:00
Markus Lehtonen	e7f87de6df	nfd-master: retry node updates indefinitely Treat node updates like a reconciliation loop. Keep trying on node update as long as it fails. Node update permafailing likely indicates a bug in the nfd code (there should be no reason for it to fail forever) and it's better to clearly see it in the logs/metrics rather than giving up after a few retries.	2024-03-18 18:14:24 +02:00
Kubernetes Prow Robot	4790962123	Merge pull request #1595 from marquiz/devel/master-check-node-existence nfd-master: check if node exists before trying update	2024-03-18 04:19:57 -07:00
Kubernetes Prow Robot	797fada92e	Merge pull request #1585 from kannon92/add-swap-support add swap support in nfd	2024-03-18 04:19:48 -07:00
Carlos Eduardo Arango Gutierrez	06c4733bc5	Add FeatureGate framework to handle new features Code inspired on https://github.com/kubernetes/component-base/tree/master/featuregate Signed-off-by: Carlos Eduardo Arango Gutierrez <eduardoa@nvidia.com>	2024-03-15 19:11:32 +01:00
Oleg Zhurakivskyy	c662265a47	topology-updater: Add owner reference to NRT object Signed-off-by: Oleg Zhurakivskyy <oleg.zhurakivskyy@intel.com>	2024-03-15 16:36:27 +02:00
Kubernetes Prow Robot	52d4337004	Merge pull request #1615 from marquiz/devel/master-mem-leak nfd-master: fix memory leak in nfd api-controller	2024-03-14 08:21:33 -07:00
Carlos Eduardo Arango Gutierrez	69dbfdfbc0	Use close to signal stop channedl in worker and topology-updater Fix stop channel management on Worker and T-updater in case of multiple callers Signed-off-by: Carlos Eduardo Arango Gutierrez <eduardoa@nvidia.com>	2024-03-14 15:28:39 +01:00
Markus Lehtonen	70fd3757c4	nfd-master: fix memory leak in nfd api-controller Fixes a memory leak that happened when stopping (and then re-starting) the nfd api controller. The stop channel was not used properly which caused the underlying informer to keep on running.	2024-03-14 15:39:10 +02:00
Kubernetes Prow Robot	b2ba043fae	Merge pull request #1605 from ArangoGutierrez/codegen Update generate scripts to use latest code_gen functions	2024-03-12 05:18:25 -07:00
Kubernetes Prow Robot	ebd19fe692	Merge pull request #1590 from marquiz/devel/validation-assert apis/nfd/validate: use testify/assert for checking test results	2024-03-12 04:44:09 -07:00
Carlos Eduardo Arango Gutierrez	dd24e8bc97	Update generate scripts to use latest code_gen functions Rewrite the generate.sh into update_codegen inspired in k8s.io/code-generator documentation. Signed-off-by: Carlos Eduardo Arango Gutierrez <eduardoa@nvidia.com>	2024-03-12 11:35:47 +01:00
Markus Lehtonen	a562a6188a	Update auto-generated code	2024-03-11 12:18:32 +02:00
Markus Lehtonen	a7bd22a75b	nfd-master: check if node exists before trying update Make the node-updater-pool worker fail fast (and not retry updates) if a node does not exist.	2024-02-20 11:04:46 +02:00
Kevin Hannon	187f65f94e	Add swap support in nfd	2024-02-19 10:20:56 -05:00
Markus Lehtonen	044fd4a3fd	nfd-master: log errors on node update retries	2024-02-16 15:51:04 +02:00
Markus Lehtonen	1f3b9ccb97	apis/nfd/validate: use testify/assert for checking test results	2024-02-16 10:05:16 +02:00
Markus Lehtonen	2382c34697	nfd-master: fix node status patching Correctly patch the "status" subresource. This got broken when refactoring the code in `7a050e7cf9` and wasn't even catched by the unit tests as the fake kubernetes client doesn't handle subresources as the real apiserver does.	2024-01-26 22:00:13 +02:00
Markus Lehtonen	8a6a731eb0	Drop pkg/apihelper The code is now unused.	2024-01-26 18:50:31 +02:00
Kubernetes Prow Robot	33858b7502	Merge pull request #1567 from marquiz/devel/apihelper-refactor-3 topology-updater: ditch apihelper	2024-01-26 17:07:13 +01:00
Markus Lehtonen	7a050e7cf9	nfd-master: ditch apihelper Implement some of frequently used helper functions inpackage. This patch also contains big changes to the nfd-master unit tests. Much of this is about migrating from the mocked apihelper interface to fake kubernetes client that provides a bit more apiserver'ish functionality. At the same time there is quite a bit of renaming in the tests, shortening and unifying naming and getting rid of the extensive usage of "mock" everywhere.	2024-01-26 16:09:22 +02:00
Markus Lehtonen	c581a25a39	topology-updater: ditch apihelper Stop using pkg/apihelper for accessing the Kubernetes API. Modify unit tests to use the fake kubernetes client instead of mocked apihelper interface.	2024-01-25 22:15:20 +02:00
Markus Lehtonen	53003cbf69	pkg/utils: move JsonPatch from pkg/apihelper	2024-01-25 17:23:14 +02:00
Markus Lehtonen	2326459d05	topology-updater: get topology api client directly Stop using apihelper for getting the noderesourcetopology-api client.	2024-01-25 16:33:34 +02:00
Markus Lehtonen	acf815fb10	pkg/utils: move GetKubeconfig from pkg/apihelper here This change is part of an effort to remove the pkg/apihelper package. GetKubeconfig is useful helper functionality shared accross the codebase so move it into a "safe" location.	2024-01-24 16:10:02 +02:00
Markus Lehtonen	57b7a3c6a8	Wrap nested errors	2024-01-22 22:45:15 +02:00
Markus Lehtonen	b452ab6a5c	topology-updater: initialize properly with -no-publish We need to parse kubeconfig (and initialize the apihelper) even with -no-publish as the PodResourcesScanner accesses the k8s API even if we're not publishing/updating NRTs.	2024-01-22 14:15:12 +02:00
Kubernetes Prow Robot	3667a4d073	Merge pull request #1537 from ozhuraki/apis-nfd-test apis/nfd: Trivial typo fix in tests	2024-01-19 15:25:41 +01:00
Markus Lehtonen	58ae81804c	go.mod: update dependencies	2024-01-15 21:29:32 +02:00
Oleg Zhurakivskyy	eec05e1c7a	apis/nfd: Trivial typo fix in tests Signed-off-by: Oleg Zhurakivskyy <oleg.zhurakivskyy@intel.com>	2024-01-15 18:06:58 +02:00
Markus Lehtonen	a053efda64	nfd-master: run a separate gRPC health server This patch separates the gRPC health server from the deprecated gRPC server (disabled by default, replaced by the NodeFeature CRD API) used for node labeling requests. The new health server runs on hardcoded TCP port number 8082. The main motivation for this change is to make the Kubernetes' built-in gRPC liveness probes to function if TLS is enabled (as they don't support TLS). The health server itself is a naive implementation (as it was before), basically only checking that nfd-master has started and hasn't crashed. The patch adds a TODO note to improve the functionality.	2024-01-04 13:58:26 +02:00
Carlos Eduardo Arango Gutierrez	57b6035b71	Add kubectl-nfd kubectl-nfd is a kubectl plugin for debbuging NodeFeatureRules Signed-off-by: Carlos Eduardo Arango Gutierrez <eduardoa@nvidia.com>	2023-12-21 16:00:19 +01:00
Markus Lehtonen	97bf841140	apis/nfd: split rule processing into a separate package This patch tidies up the nfdv1alpha1 API package by refactoring out the implementation of (NodeFeature)Rule evaluation into a separate package.	2023-12-20 12:52:15 +02:00
Gyuho Lee	ed0418b81c	chore(nfd-worker): fix minor typo in wrong label value format error Signed-off-by: Gyuho Lee <gyuho@lepton.ai>	2023-12-19 02:29:37 +08:00

1 2 3 4 5 ...

432 commits