This is a question many have asked but the answer is never right for all scenarios. The question is coming up again with my EventScavenger system but the answer is not obvious. There are benefits to both polling and not polling (usually by something that raises events). This is really an architectural question and I’m kind of just blamestorming by myself (blaming myself among other things).
A quick overview of the definitions (in my context of my scenario):
What is polling?
Usually this is an action where something running locally periodically access something remotely and then gather data, do some work (store data) and then go to some sleep or waiting state again until a next scheduled time.
What is the alternative (not polling)?
Usually this involves something at the source pushing the data you need locally to you (or the data storage place where you can access it). You also must have something that ‘subscribe’ to this event to capture it.
What are the benefits of each/each approach?
Polling:
- Not active the whole time – i.e. open connections over network, resource handles etc. It spends time ‘waiting’ until next scheduled polling event.
- Don’t require remote software to be installed (something that subscribes to the events)
- No ‘missed’ events to worry about. On each poll you gather all relevant data since previous polling anyway.
- Easier to administrate as it is usually a central component/service.
Not polling (Events):
- Data gets transferred/stored ‘as it happens’.
- No big batches of data send over network – each event’s data is sent on its own.
- No network traffic while there is no data to report on
- Number of remote locations not (so much) limited to what a single ‘poller’ could handle.
Unfortunately there is no golden ‘middle way’. Its either the one or the other approach (per resource you want to access). The current approach for EventScavenger is plain old polling (each event log is done on a separate thread. This works well up to a point until (1) the number of threads become too many, (2) the number of events per ‘poll’ get so many that the particular thread cannot process things quickly enough. To overcome the these problems you can install an instance (collector as named in the EventScavenger context) on the source machine. That can work if you access to the machine (are allowed to install) but then you might as well ask why not use the event driver approach?
The flip side is with subscribing to events is that you must have something installed ‘locally’ on the resource (eventlog in this case) to access the data. That requires something installed and/or running extra there. Also, if for any reason the subscriber is down/busy or something the events raised would be missed with no way to get them back again afterwards.
I’ve been looking at the EventLog.EntryWritten event and also the whole new System.Diagnostics.Eventing.Reader.EventLogWatcher class. Both only works for the local event logs and I’m having issues with the newer EventRecord class that does not give me the proper full description message (always return a null). Perhaps I have to look at an additional component that runs alongside the existing ‘collectors’ services to gather ‘problem child’ event logs – but on the actual source machine and just for raised events. This could help other people but I’m having a situation where I cannot install something on the sources of the logs that must be imported since they are ‘off-limits’ (Domain controllers and Windows Core installs)
Why can’t life be easy? 😉