Internet commentary
Chips with everything [1]
Keywords: Internet, Chip scale packaging
I wrote an article entitled Spam, Spam, Spam in our sister journal, Soldering and Surface Mount Technology, Volume 15, Issue 2. Guess what it was about?Yes, of course, spam or junk e-mail. One of the most important features of the Internet is e-mailing. In fact, it is a mainstay of business today and has rendered the fax almost obsolete and, well, the telex machine, clattering in a corner, is really past history. Spam is becoming an ever- increasing nuisance and I explained in that article that I used to lose the equivalent of a working day per year just identifying and deleting spam. I am told that time is money,so the inconsiderate blighters who sully the Internet at almost no cost to themselves cost us all money. Nowadays prologue is therefore an updated summary of that article and describes how spam can be all, but eliminated. However,before I start, if you are on a corporate network, do not install any anti-spam software without having discussed it with your system administrator, first, for fear that you upset his delicately balanced applecart. If you are using a computer with the Internet connection individual to it, then do not hesitate,there is no problem in installing what you like.
When I first started to become bothered by spam, I resorted to filters in my e-mail client; you know the thing, if the Subject contains Viagra, send to Trash. This became inadequate many years back, even with tens of complex filters. The criminals (and this is not too harsh a word, because they steal my time and the money it costs to download their trash) who perpetrate the felony of imposing unwanted junk on us started to get smart and took measures so that ordinary filters would not work. Then came pattern filters which did not work entirely on the word, but on how the messages were set out. These worked for a while, but even then they became imperfect, as the messages took on other patterns and key words were split by an "invisible space", which upset ordinary word recognition.
At this stage, I resorted to a pre-filter application, called Mailwasher,combined with filters to sort out the messages into the right mailboxes. I had high hopes for this, but it failed me in the long run. Before downloading the messages, into an e-mail client, it would list the unknown senders and subjects. If any message on the ISP server were from a known spam source, it would delete it without listing it. You quickly went through the list and marked anything that appeared to be spam. In a treatment phase, it would record the name of the spam senders and delete the messages from your ISP server, while opening your default e-mail client which would then download only your wanted messages. As such, it was faster than deleting from the client. On request, it would also"bounce" the spam e-mails back to the sender, pretending that your e-mail address did not exist. Unfortunately, this took a lot of extra time. Why did this fall down? Well, firstly, spammers rarely use the same address twice. This resulted in an enormous file of many hundreds or even thousands of spammers. The worst thing was if you accidentally checked a wanted message as spam, which meant that you never saw a message from the sender again. Of course, if you realised this, you could remove his address from the blacklist or even put it into the white list, but the risk is real, perhaps one in a thousand with careful use, but real. At first, I wasted a lot of time using the bouncing feature, thinking that if I systematically did so, the name would eventually be removed from the lists of millions of addresses that are sold to unsuspecting spammers. Even after 2 years of bouncing, the fury of incoming spam only increased. This was not the answer.
Then I chanced on the name William S. Yerazunis, a Computer Researcher at the MIT. He developed a system called the Controllable Regex Mutilator or CRM. (see http://crm114.sourceforge.net/). This used a technique called the Bayes' theorem which can be stated as providing a way to apply quantitative reasoning to what we commonly think of as scientific. When several alternative hypotheses are competing for credibility,we test them by deducing the consequences of each one, then conducting experimental tests to observe whether or not those consequences happen. If a hypothesis foretells that something should occur and it does, our belief in the veracity of the hypothesis is stronger. Conversely, if the experiment does not verify the prediction, it would weaken our confidence. So how does the Bayesian notion apply to spam elimination?
Each message (including the header and subject) is scanned by the system and a number of "unique" keywords are extracted. Each is analysed and a probability is calculated as to whether it is spam or wanted. By calculating the overall probability for the whole group of keywords fitting into the category of spam,or not, so the message can be accepted or rejected. At this stage, there is a possibility of manual correction, so the system can actually learn which keywords are used in the typical mail you, as an individual, receive and which you do not use. To take an extreme example, the word "sexy" may be used extensively in spam, but a gynaecologist may use the word "sex" professionally quite frequently. The analysis may allocate a spam probability of, say, 0.9988 to the first word, leaving a 0.0012 probability that it is wanted. The second word may give probabilities of, say, 0.4000 and 0.6000, respectively. Of course,there are other keywords in each message. For example, if the message contained"uterus" or "ovaries", then this would weigh heavily in favour of a wanted message, but if it contained salacious obscenities, then it would almost certainly be spam. Common words are excluded as keywords.
Unfortunately, Yerazunis' developments are available only for Linux and Unix platforms, so this is useless for those working with Windows. So I did a search and came up with a Windows' Bayesian spam filter, called POPFile. This normally uses 20 keywords in each message. On the analysis of these, it allocates each message to a "bucket" which may be named as you wish. You can have any number of buckets, but six-ten would be a reasonable practical maximum, under most circumstances, although it is possible to work with just two which you may wish to name "spam" and "wanted", for example. To start with, the identification of wanted and spam messages is only about 50 per cent correct. However, there is a bias towards wanted ones. This means that you are much more likely to find a spam message in the wanted bucket than a wanted message in the spam bucket. As time goes on, so the accuracy increases, stupendously. By the time the system has analysed just 2,000 words in each bucket, of which it may have selected 500 or 600 "unique" keywords, the accuracy increases to something like 95 per cent or more. At no later than this stage, "false spam" or wanted messages being classed as spam is negligible. At about 2,000 "unique" words, per bucket, the precision is typically 99 per cent or better. I installed POPFile several weeks ago and have received about 100 or so messages per day in that time. I initially set up nine buckets, but I quickly learned that it is possibly better not to have buckets receiving little mail. I have now reduced it to six buckets. With these, I have had over 99 per cent accuracy (it varies daily between about 98.9 and 99.3 per cent). Typically, the average daily errors are one spam message arriving in a wanted inbox. I have not had a wanted message appearing in the spam box for 2 or 3 weeks and it is also equally rare for a wanted message to go to the wrong wanted inbox, even with overlapping subject matters.
So, how does POPFile interface with the e-mail client? There are two options. The universal one is that it can insert the bucket name between the square brackets in the Subject line of the e-mail. This can then be used to filter them in the e-mail client. The second method is to add an extra line in the header in the form X-Text- Classification: <bucketname>. The latter is much cleaner, as it is invisible and it would not appear in the Subject line if you reply to a message.
How does it work, in practice? The one-word answer is invisibly. You open your e-mail client, as you would do normally, elect to receive your messages and you see them roll in, allocated to their correct mailboxes (if you use multiple buckets or filtering) just as you would do without POPFile, except that all the spam messages are separated and placed where you choose to put them. At the moment, I filter them to the Trash box, where I can double check against false spam before deleting, but when I am absolutely confident that there is no false spam, I will delete it beforehand, using the appropriate filter.
Some final points about POPFile: it will handle some foreign languages in place of English, but it will need more specific training if you want to handle two or more languages simultaneously. I have received some spams and wanted messages in French since I started. The first spam totally perplexed it, but I just reclassified it to the spam bucket. The second one (from the same source)went there of its own accord, so it must have recognised something in common with the first one. The first wanted message was long and technical and it recognised a few technical words in common with English, so it actually classified it correctly, to my surprise. I believe that, if I were to receive much mail in a foreign language, I would add many of the common words to the list of "Ignored" words, in the Advanced tab of the software interface (these are words like "the", "is", "have" and so on). Otherwise, I think it would cope admirably.
Will the spammers be able to cause Bayesian filters not to recognise their masterpieces? The answer is a mitigated yes, in the short term. The spammers always try to keep one step ahead of the competition and they are well aware that Bayesian filtering is defeating their ends. The most successful way (over half the classification errors, in my case) seems to be to have a very short,anodyne, message, using words such as Jim may send to Joe, without tell-tale key words, followed by a hyperlink to a Web site, which is the real message. However, even this will work only once, if either the same Web site URL or the same anodyne message is used twice. There are many other techniques, to invisibly split words with a useless tag, for example, or to put a mass of thousands of acceptable words in white on white, so that there is less chance of the unacceptable words being found. Actually, POPFile now assumes a strong probability that coloured text on the same-coloured background will be spam,because there is little reason for hiding such text in a genuine e-mail. Anyone using the system will soon master the finesses of the spammers and the anti-spammers, because it is constantly being updated to counter new techniques.
You can find more details at http://popfile.sourceforge.net/but do not be put off by the appearance of the Home Page: it is utilitarian and unpretty. Here, you will find this is an Open Source application, working under a Perl interpreter and, oh joy!, not only free of charge, but free of advertisements. Installation into a Windows operating system is fairly straightforward, provided you carefully follow the instructions. It will also work under other OSs, such as Mac, Solaris and Linux. Configuring POPFile for your e-mail client is also straightforward, again if you follow the instructions, except for Netscape/Mozilla (contact me if you want to find the secret for this type of client, which actually marries very well into it).
This prologue has been somewhat longer than usual, but I believe the subject is important. For my review section, I will follow on from an excellent new book I have just finished reading, Chip Scale Packaging for Modern Electronicsby Joseph Fjelstad et al. (ISBN 0 901150 43 6).
This describes an ultra-thin stacked die chip scale packaging (CSP), used in this case for wireless memories. There is little new in this technique; it is used by many chip makers, especially for mobile phones. What is advanced is that Intel have achieved a five-die-high stack in a package height of 0.8mm,apparently by thinning the dies. The utility for this, say, PDAs is obvious.
This is a useful paper entitled Flip Chip, CSP and WLP Technologies: A Reliability Perspective. It is published by a manufacturer of underfill dispensers and, not surprisingly, extols the virtues of this technique! It also reviews the historical evolution of the use of flip chips and their incorporation into CSPs.
One of the key components in modern electronics is the programmable non-volatile memory, which holds the firmware for myriad devices. Xilinx claim to make the smallest such components in the world. I guess it has no mean feat to make a package with a pin-out of 56 taking a space of just 6×6 mm,but it would not be child's play, either, to assemble it on the HDIS board of a mobile phone or a miniature camcorder. Our technology is advancing by leaps and bounds!
Here is a CSP without a chip! In fact, it is a thin film filter circuit packaged into a CSP-style assembly. Of course, this notion is not new; we have had, for example, dual-in-line resistor arrays in similar packaging to ICs for many decades, so why not in packages that look like CSPs? Unfortunately, I suspect that this product has not yet been fabricated, because the manufacturer has stated that the layout/design is pending! Oh, and by the way, there is an amusing error: the page states that the resistors are 100 W, rather than 100Ω! Maybe this is the first CSP heating element for cooking chips!
As dies shrink, according to a corollary of Moore's Law, for a given functionality, so do CSPs. Ball sizes of 0.25-0.30 mm on a rectilinear pitch of 0.50 mm are increasingly common. How do you test or burn in such devices without damaging the balls, making subsequent use of the devices less reliable? This paper makes some suggestions, although the ideal answer may not exist.
Large pin-out numbers increase the complexity of a CSP and its interconnections. Would it not be nice to have ICs with just a signal and ground connection? In effect, this would look more like a resistor, but this page suggests that a number of functions may be done with just that. The same connection is used for both power and signal in and out. This ingenious idea even allows communications between different CSPs with similar physical forms.
It is not often that the IVF offers a good and useful freebie; this is one. It is a series of pages describing CSPs, their advantages and disadvantages and a whole mountain of technical data. This is mandatory reading for those interested in the subject: do not miss it!
This is a good paper discussing the pros and cons of CSP for hi-rel applications. Even though it is a few years old, it is still well worth the read. Certainly, there has been an evolution since 1998, but nothing sufficiently mind-boggling that renders this paper out-of-date. Even though the site is hosted by NASA, the contents are applicable equally to commercial applications.
There is an enormous wealth of information on all aspects of CSP technology on the Internet. What I have offered here is just the smallest tip of the iceberg. If you Google with "chip scale package", you will get links to nearly 8,000 sites, of which these are just a small selection. I believe that this subject is becoming of increasing importance and I forecast that CSPs will become the mainstream semiconductor packaging within a few years, even for those applications where weight or size is not important. For ceramic substrates, the flip chip will become even more dominant.
Brian EllisCyprus
Note
- 1.
Title of a play by Arnold Wesker, 1962. This phrase was falsely attributed to Lord Randolph Churchill in a political speech in Blackpool in 1884. What he actually said was, "He [Gladstone] told them that he would give them and all other subjects of the Queen much legislation, great prosperity, and universal peace, and he has given them nothing, but chips. Chips to the faithful allies in Afghanistan, chips to the trusting native races of South Africa, chips to the Egyptian fellah, chips to the British farmer, chips to the manufacturer and the artisan, chips to the agricultural labourer, chips to the House of Commons itself." (For those used to other idioms, "chips" in the UK are a usually thick and soggy version of what is known elsewhere as French fries, as well as semiconductor dies.).
