Article navigation
  • Spam, Spam, Spam[1]

Keywords: Internet, E-mail, Surface mount technology

I have written many times before on this subject but it is the one that evolves continuously, and hence merits further discussion. One of the most important features of the Internet is e-mailing. In fact, it is a mainstay of business today and has rendered the fax almost obsolete and, well, the telex machine, clattering in a corner, is really past history.

E-mailing has a number of down sides. Not the least of these is spam(theoretically, I should write this in inverted commas, but I consider the word has now really entered into our language, thanks to a BBC comedy show). For the uninitiated, spam is a collective term for those messages, usually advertising ink-jet ink, pharmaceuticals, university degrees to buy, sexual services of all flavours, how to make a fortune at home and such-like junk mail. Just to show how bad it has become, out of an average 70 e-mail messages per day I receive,about 40 percent are spam on weekdays and 60 per cent during the weekends. How can we deal with this plague? Deal with it, we must for, if we handle it manually, it takes about 3s to identify and delete each one. In my case, I estimate I receive about 12,000 spam messages per year, so it would take me more than a working day in this time just to identify and erase them. Hopefully, I would not miss any, nor, worse, would I accidentally erase a wanted message, errare humanus est.

Initially, I resorted to filters in my e-mail client; you know the thing,if the subject contains Viagra, send to trash. This became inadequate many years ago, even with tens of complex filters. The criminals (and this is not too harsh a word, because they steal my time and the money it costs to download their trash) who perpetrate the felony of imposing unwanted junk on us started to get smart and took measures so that ordinary filters would not work. Then came pattern filters, which did not work entirely on the word but how the messages were set out. These worked for a while, but even they became imperfect,as the messages took on other patterns and key words were split by an "invisible space", which upset ordinary word recognition.

At this stage, I resorted to a pre-filter application, called Mailwasher,combined with filters to sort out the messages into the right mailboxes. I had high hopes for this, but it failed me in the long run. It worked on a different principle, along with a partial learning process. Before downloading the messages, in an e-mail client, it would list the unknown senders and subjects. If any message on the ISP server were from a known spam source, it would delete it without listing it. You quickly went through the list and marked any that appeared to be spam. In a treatment phase, it would record the name of the spam senders and delete the messages, while opening your default e-mail client which would then download only your wanted messages. As such, it was faster than deleting from the client. On request, it would also "bounce" the spam e-mails back to the sender, pretending that your e-mail address did not exist. Unfortunately, this took a lot of extra time. Why did this fall down? Well,firstly, spammers rarely use the same address twice. This resulted in an enormous file of many hundreds or even thousands of spammers. There were means of reducing the list. For example, if you saw there were many addresses like john17392@isp.com where only the number changed each time, you could change it to john*@isp.com where the wildcard took care of future numbers and you could tell the system that a regular real correspondent using the same ISP with the address johnsmith should not be bumped off. You could also tell the system to delist any spammers' addresses that had not been used for over a given length of time. The worst thing was if you accidentally checked a wanted message as spam; that meant that you never saw a message from the sender again. Of course, if you realised this, you could remove his address from the blacklist or even put it into the white list, but the risk is real, perhaps one in a thousand with careful use, but real. At first, I wasted a lot of time using the bouncing feature, thinking that if I systematically did so, the name would eventually be removed from the lists of millions of addresses that are sold to unsuspecting spammers. Even after 2 years of bouncing, the fury of incoming spam only increased. This was not the answer.

Then I chanced on the name William S. Yerazunis, a computer researcher at the MIT. He developed a system called the controllable regex mutilator (CRM) (see http://crm114.sourceforge.net/). This used a technique called the Bayes' theorem, which can be stated as providing a way to apply quantitative reasoning to what we commonly think of as scientific. When several alternative hypotheses are competing for credibility, we test them by deducing the consequences of each one, then conducting experimental tests to observe whether those consequences happen or not. If a hypothesis foretells that something should occur and it does, our belief in the veracity of the hypothesis is stronger. Conversely, if the experiment does not verify the prediction, it would weaken our confidence. So, how does the Bayesian notion apply to spam elimination?

What happens is that each message is scanned (including the header and subject) and a number of "unique" keywords are extracted. Each is analysed and a probability is calculated as to whether it is spam or wanted. By calculating the overall probability for the whole group of keywords fitting into the category of spam, or not, the message can be accepted or rejected. At this stage, there is a possibility of manual correction, so the system can actually learn which keywords are used in the typical mail you, as an individual, receive and which you do not use. To take an extreme example, the word "sexy" may be used extensively in spam, but a gynaecologist may use the word "sex" professionally quite frequently. The analysis may allocate a spam probability of, say, 0.9988 to the first word, leaving a 0.0012 probability that it is wanted. The second word may give probabilities of, say, 0.4000 and 0.6000, respectively. Of course,there are other keywords in each message. For example, if the message contained"uterus" or "ovaries", then this would weigh heavily in favour of a wanted message, but if it also contained salacious obscenities, then it would almost certainly be spam. Common words are excluded as keywords.

Unfortunately, Yerazunis' developments are available only for Linux and Unix platforms, so this is useless for those working with Windows. So, I did a search and came up with Windows' Bayesian spam filter, called POPFile. Let it be said,here and now, that it is not a pretty interface to look at, but it does seem to work. There is a choice of "skin", and the "tinygrey" is possibly the most user-friendly, although the colour scheme is not attractive. This normally uses 20 keywords in each message. On the analysis of these, it allocates each message to a "bucket", which may be named as you wish. You can have any number of buckets, but 6 to 10 would be a reasonable practical maximum, under most circumstances, although it is possible to work with just two which you may wish to name "spam" and "wanted", for example. To start with, the identification of wanted and spam messages is only about 50 percent correct. However, there is a bias towards wanted ones. This means that you are much more likely to find a spam message in the wanted bucket than a wanted message in the spam bucket. As time goes on, the accuracy increases, stupendously. By the time the system has analysed just 2,000 words in each bucket, of which it may have selected 500 or 600 "unique" keywords, the accuracy increases to something like 95 percent or more. By this stage, "false spam" or wanted messages being classed as spam is negligible. At about 5,000 analysed words and 2,000 "unique" words, per bucket,the precision is typically 99 percent or better. I installed POPFile just 2 weeks ago and have received about 1,000 messages in that time. I have set up nine buckets. The least-used one has analysed 1,010 words with 509 "unique"words and the most used wanted one 8,177 and 1,089, respectively. The spam bucket has analysed 8,348 and 2,435, respectively: note the greater proportion of "unique" words. I reset the statistics before 228 messages and there have been three wrong classifications in that period. In no case there was either false spam or wrongly classified spam was found. The three "errors" were in the choice between two of the lesser-used wanted buckets, both work-oriented with an overlap of terminology. This 98.68 percent success is, therefore, convincing,for a short term trial. I am sure that if I had set up just two buckets, the success rate would have been 100 percent.

So, how does POPFile interface with the e-mail client? There are two options. The universal one is that it can insert the bucket name between square brackets in the Subject line of the e- mail. This can then be used to filter them in the e-mail client. The second method is to add an extra line in the header in the form X-Text- Classification:<bucketname>. The latter is much cleaner, as it is invisible and would not appear in the subject line if you reply to a message. My first trials were with the Netscape 7.01 messenger and I could not find a means of adding a header line into the filter, so I started with the first method, inserting the bucket name into the Subject. I quickly decided that this was a royal pain. According to the documentation, the header method was usable by Eudora. I have heard that many Eudora users were more than enthusiastic about their e-mail client and, as I had not tried it for several years, I downloaded the free version. I persevered with it for 13 days and found that it did, indeed, have some excellent ideas. But I also found it slow and severely lacking in "user- friendliness" and even some basic features. Because I was thinking that I may have been missing something, I joined an internet forum and asked some questions regarding how to do this or that. To my surprise,instead of receiving an answer or being politely told that the feature was not available, the forum moderator flamed me with a long rant, presumably because she thought I was subtly attacking her beloved software as incomplete, where I was only seeking a shortcut to go from one unread message to the next! This rather sickened me, so I had another, deeper look at Netscape and found a poorly documented way of using the header method. For those who wish to know how it is done, go into message filters and start a new one. In the window that comes up,click on the arrow next to "subject". At the bottom of the list, click on"customize". A dialogue box comes up and you can add X-Text- classificationto the list. This, then allows you to add a filter or filters for each bucket. A small documented change to the server settings and it becomes operational.

How does it work, in practice? The one-word answer is invisible. You open your e-mail client, as you would do normally, elect to receive your messages and you see them roll in, allocated to their correct mailboxes (if you use multiple buckets or use filtering) just as you would do without POPFile, except that all the spam messages are separated and placed where you choose to put them. At the moment, I filter them to the Trash box, where I can double check against false spam before deleting, but when I am absolutely confident there is no false spam. I will delete it beforehand, using the appropriate filter (incidentally, not possible in Eudora).

Some final points about POPFile: it will handle some foreign languages in place of English but it will need more specific training if you need to handle two or more languages simultaneously. I have received two spams and one wanted message in French since I started. The first spam totally perplexed it, but I just reclassified it to the spam bucket. The second one (from the same source)went there of its own accord, so it must have recognised something in common with the first one. The wanted message was long and technical and it recognised a few technical words in common with English, so it actually classified it correctly, to my surprise. I believe that, if I were to receive more mails in a foreign language, I would add many of the common words to the list of "ignored"words, in the advanced tab of the software interface (these are words like"the", "is", "have" and so on). Otherwise, I think it would cope admirably. Will the spammers be able to cause Bayesian filters not to recognise their masterpieces? The answer is a mitigated yes, in the short term. For example,they try to do this by hiding them in graphics or breaking up key words into two or three non-syllabic fragments. For example, they may break up "mortgage" (a common spam word) into "mo.rtg.age" where the breaks are a character unrecognised in HTML, usually a special space or something like </d> or</h>, so they close up. The word looks normal to the eye but seems like three words to the Bayesian analyser, because tags identify the end of a word. It will learn "rtg" as a unique keyword, which will be given a high probability in spam, so the next message using this technique will be recognised for what it is. Alternatively, the word is encoded and decoding occurs in the e-mail client. I have little doubt that they will try more sophisticated techniques as time goes on and we all know that it is sometimes difficult to keep one step ahead. The most dangerous method of circumventing these techniques is to keep the message as short and banal as possible. For example, I received a spam, about a week ago, that simply said, subject: "requested information" and the message:"The information you requested can be found at http://www....". POPFile was very undecided about what to do with it: it gave pretty well equal chances between spam and wanted. It played safe and told me that it was wanted, but put it into a rarely-used bucket for odds and ends. I could hardly fault it for that, as there were really only two usable keywords, although it did find some more in the header and alias routing. Finally, does it have any effect on speeds, and resources? Undoubtedly, yes. There is a small permanent use of memory for the Perl interpreter, but this would go into the swapfile when not being used, if it became critical. The process time for the incoming messages is certainly less than the download time, using a 56kbit/s modem. To me, it is, therefore, totally invisible. Having said that, those downloading their e-mail through a broadband connection may see it coming in slightly slower. If you wish to know more or try it out, see http://popfile.sourceforge.net/. It will cost you nothing, as it is a open source freeware. It does not upset anti-virus detectors or firewalls, although you do have to give permission for it to pass the latter on a port you designate.

This prologue has been somewhat longer than usual, but I believe the subject is important. As usual, I will review an internet site relevant to our industry,but this section will be shorter than you may be used to seeing – sorry!

http://www.smtinfocus.com/

I came across this site by accident, one day. It is a web portal that, as its URL implies, is devoted to surface mount technology, or rather indexes other pages that do. Of course, like all portals, it lives from advertising and this is something "up with which, we must put" (as some English grammar purists insist we should say: personally, I think "we must put up with" is far better!). Browsing down the centre pane of the page brings us some information on what is happening in our world and some links to the trade organisations in our industry. In the left hand column, there is a menu, on which I will concentrate.

The first item leads us to a question and answer forum. I am not particularly keen on the way it works. There are better ones in terms of layout and usability. However, that does not detract from the quality of the replies,mostly made by a small core of individuals. I scanned through a few of them and found that the help given was reasonable. Notwithstanding, one cannot help but feel that the large number of subscribers to the IPC TechNet netlist would offer a better chance of hitting on a wider range of opinions, especially on controversial subjects. The number of subjects per unit time is also much smaller than TechNet. This brings me a question: Is it better to have many forums and netlists "competing" with one another or just using one big one?Obviously, in this case, the aim is to have users read as many of the advertisements as possible and click on them, but there are no ads on the forum page. I do not know.

The next section is a buyers' guide. This is good, assuming that each section is as complete as possible. One assumes that the companies advertising therein pay for their mention. The sections include: screen printers; SMD placement;Reflow ovens; 2nd hand equipment and soldering materials. In the case of equipment-oriented pages, each manufacturer gives a list, with very basic specifications of each of his machines. For example, under SMD placement, there are 122 machines listed from 18 manufacturers, claiming speeds up to 96,000 chips per hour (as at the time of writing). Soldering materials have merited a slightly different treatment because the page would otherwise have been 10km long. Each of the 17 manufacturers has a hyperlink to his home page with an indication of the generic types of products on offer. 2nd hand equipment(variously called "pre-owned SMT equipment", depending on the page, a rose is a rose!) is just simply a couple of lists of known dealers in used machines.

This section is, followed by the one on information to help people fine-tune their processes, divided into: SMT failure library; SMT adhesive dispensing;solder paste printing; SMD placement and reflow soldering. The first of these opens up to a page of thumbnail photographs. Clicking on one of them will produce an enlargement, with a description of the cause, effect and tips for a cure. The rest of the pages are fairly long and comprehensive treatises on each of the subjects, at a medium technical level. They are, on the whole, accurate but there are a few sins of omission and commission on each of them.

The next part of the menu is, rightfully, headed "miscellaneous". The first item is a list allowing you to look at about 35 technical papers on many related subjects. These are mostly in PDF format. The quality of the individual papers is variable, as may be expected. Some of them are important contributions. The next page is devoted to lead- free issues and lists the legislative situation,some technical papers and links to Web sites devoted to this type of technology. Then comes "Dave's bookshelf", subdivided into three pages. The first is what David Fish considers essential reading; the second is entitled "to learn more about SMT", with a list of "better books" and a second list of "books that are not worth the money"! The last page is "to learn more about productivity". Then comes a page advertising free magazines, not all on relevant subjects. After that is a calendar of about a dozen global trade shows per year (with the 2002 ones still advertised in February 2003!). A page of very simple definitions follows: could be more complete and less terse, to make it better. There is then a list of industry hyperlinks in a single page: again, rather on the short side.

The last section is devoted to various things about the publisher of the portal, forms to receive a newsletter, receive details about advertising and sponsoring, linking and so-on.

On the whole, my judgement is that the site is good and useful, although lacking information on a number of important issues. Notably, I saw nothing related to the protection of the environment or health and safety, even though our industry is noted as being, both directly and indirectly, a polluter. I include lead-free soldering in this statement because it has nothing to do with these subjects, even though it is often falsely presented as an environmental issue. In terms of appearance, I believe it could be brightened up a little;white pages with wide black borders give a rather funereal impression.

An excellent site, well worth the visit!

Because this Commentary is somewhat longer than usual and the subjects treated do not specifically benefit from one, I have not put in a screenshot.

Brian EllisCyprusb_ellis@protonique.com

Note

  • 1.

    Use of the term "spam" was adopted as a result of the Monty Python skit in which a group of Vikings sang a chorus of "Spam, Spam, Spam... " in a crescendo,drowning out other conversation. The analogy applied because unsolicited commercial e-mails were drowning out normal "conversation" on the Internet. It has nothing to do with the canned luncheon meat of the same name that is no longer available, but became a bit of a joke during the second world war in the UK. Lord Woolton, the then minister of food, also perpetrated a number of other doubtful jokes on the unsuspecting but hungry public, not the least of which was"snoek" a very inferior canned fish, related to the tuna, which was supposed to replace salmon.

 

or Create an Account

Close subscription notice
Close access options