![]() |
| September 25, 2018 | Volume 14 Issue 36 |
Manufacturing Center
Product Spotlight
Modern Applications News
Metalworking Ideas For
Today's Job Shops
Tooling and Production
Strategies for large
metalworking plants
FAULHABER GPT gearheads for miniature and micro motors deliver high power density, exceptional flexibility, and excellent cost efficiency. Designed for seamless integration with diverse motors and encoders, these compact planetary drives feature hardened stainless steel components that reliably withstand extreme torques and abrupt load changes. Offered in standard, low-noise, and high-torque variants, they ensure precise, durable performance across wide operating temperature ranges.
Learn more.
If you are having a problem with your linear guides not always staying perfectly straight during use, it may be due to a phenomenon called waving -- a problem that is particularly critical in high-precision markets such as semiconductor and LCD equipment-related applications or machine tools. Thankfully, THK has an answer.
Read the full article.
NORD DRIVE-SYSTEMS offers robust, highly configurable drive solutions designed to optimize efficiency, ensure hygiene, and reduce the TCO for automated bakery systems. From industrial-scale mixing and portioning to baking, cooling, and packaging, NORD's modular product portfolio delivers precise, reliable performance across every stage of production.
Read the full article. You may learn something even if you are not in the baking industry.
Battery-powered motor applications require careful design considerations to pair motor performance and power consumption profiles in concert with the correct battery type. This Power Electric article covers power requirements, performance considerations, and battery choices to assist you in selecting an efficient motor and a battery with the appropriate capacity. Good technical info.
Read the Power Electric technical article.
The Sinamics G210X is a new frequency converter for advanced pump, fan, and compressor applications, combining easy engineering with robust design. Integrated functions, seamless TIA Portal integration, and a web server reduce PLC effort to speed up commissioning, operation, and diagnostics. IP55 protection, 3C3 coating, and S2 system redundancy ensure reliable operation in demanding environments.
Learn more.
Curtiss-Wright's Actuation Division has expanded its Exlar line with hygienic electric actuators using FDA-approved materials and finishes. Designed for food, beverage, packaging, and pharmaceutical automation, the new GTF unit enables economical USDA, 3-A, BISSC, and EHEDG certification. Its IP69K washdown option, inverted roller screw, and compact servo-driven design deliver reliable, high-performance motion for hygienic machinery.
Learn more.
SEW-EURODRIVE is helping power one of the most ambitious bulk material handling projects in North America through its contribution to the Dune Express conveyor system, a record-setting 42-mile single-flight conveyor across the Permian Basin.
Read the full article.
Built on Copley's proven NanoPlus platform, the compact, 1.2-oz R-Series Nano servo drives withstand extreme temperatures, vibration, shock, and humidity. Available in R47 (CANopen) and R48 (EtherCAT) models, they suit space-constrained, harsh environments such as mil/aero robotics and gimbals. This commercial off-the-shelf series provides a hardened option without defense-specific development lead times.
Learn more from Copley Controls.
Kyntronics' new all-electric HyCore hybrid actuator is a compact, low-cost alternative to traditional hydraulic, pneumatic, and electro-mechanical systems. Engineered for OEMs, it delivers up to 14,726 lb of force. The self-contained design eliminates leaks and wear, providing precise control, shock-load tolerance, and high efficiency. It simplifies integration and lowers operating costs for mobile equipment, packaging, assembly automation, and more.
Learn more.
SDP/SI's best sellers aren't just popular -- they're proven. These are the motion components their customers return to time and again for precision, reliability, and unbeatable value. From belts and pulleys to gears, bearings, and couplings, each product has earned its place through consistent performance in real-world applications.
Learn more and see the full products list.
Grinding large fabrications is a classic dull, dirty, and dangerous task perfectly suited for automation, yet traditional setups require multiple costly robots. At Automate 2026, Güdel debuted a single-robot solution that utilizes two extra degrees of freedom to finish massive surfaces without complex part repositioning.
Read the full article.
Automation-Direct now offers US-manufactured RBS ball screws and nuts for precise linear motion in OEM and maintenance applications. Achieving over 90% efficiency, they feature low friction, high accuracy, and minimal wear. Available in various diameters, lengths, and leads, these precision-matched components handle axial loads with minimal backlash, providing a dependable solution that reduces maintenance and extends operational life. Great prices too.
Learn more.
Building on its established portfolio of twin-strand conveyors that are currently used in a variety of industry verticals, Bosch Rexroth is introducing the TS 7plus transfer system, which is the world's first freely configurable, fully electric conveyance solution to workpieces weighing up to 3,000 kg. TS 7plus transports material via conveyor rollers on freely configurable modular sections with lift/transverse, rotary, and positioning units, as well as stop gates. Great for automotive, battery, aerospace/defense, and more.
Learn more.
The A-123 noncontact nanoposi-tioning stage integrates a brushless motor, air bearings, and a 1-nm encoder, supporting 40-kg payloads over 750-mm travel. Its pressurized air film eliminates the friction, wear, and vibration of mechanical stages, ensuring zero particle generation. Customizable with options such as granite bases and isolation systems, it is ideal for semiconductor metrology, inspection, photonics, and more.
Learn more.
Some of the recent research activities in the area of electric motor drives for safety-critical applications (such as aerospace and nuclear power plants) are focused on looking at various fault-tolerant motor and drive topologies. After discussing different solutions, this article focuses on a miniature permanent magnet (PM) stepper motor design that provides increased redundancy.
Read this informative FAULHABER article.
By Rob Matheson, MIT
MIT computer scientists have developed a system that learns to identify objects within an image, based on a spoken description of the image. Given an image and an audio caption, the model will highlight in real time the relevant regions of the image being described.
Unlike current speech-recognition technologies, the model doesn't require manual transcriptions and annotations of the examples it's trained on. Instead, it learns words directly from recorded speech clips and objects in raw images, and associates them with one another.
The model can currently recognize only several hundred different words and object types. But the researchers hope that one day their combined speech-object recognition technique could save countless hours of manual labor and open new doors in speech and image recognition.

MIT computer scientists have developed a system that learns to identify objects within an image, based on a spoken description of the image. [Image: Christine Daniloff]
Speech-recognition systems such as Siri, for instance, require transcriptions of many thousands of hours of speech recordings. Using these data, the systems learn to map speech signals with specific words. Such an approach becomes especially problematic when, say, new terms enter our lexicon, and the systems must be retrained.
"We wanted to do speech recognition in a way that's more natural, leveraging additional signals and information that humans have the benefit of using, but that machine learning algorithms don't typically have access to. We got the idea of training a model in a manner similar to walking a child through the world and narrating what you're seeing," says David Harwath, a researcher in the Computer Science and Artificial Intelligence Laboratory (CSAIL) and the Spoken Language Systems Group. Harwath co-authored a paper describing the model that was presented at the recent European Conference on Computer Vision.
In the paper, the researchers demonstrate their model on an image of a young girl with blonde hair and blue eyes, wearing a blue dress, with a white lighthouse with a red roof in the background. The model learned to associate which pixels in the image corresponded with the words "girl," "blonde hair," "blue eyes," "blue dress," "white light house," and "red roof." When an audio caption was narrated, the model then highlighted each of those objects in the image as they were described.
One promising application is learning translations between different languages, without need of a bilingual annotator. Of the estimated 7,000 languages spoken worldwide, only 100 or so have enough transcription data for speech recognition. Consider, however, a situation where two different-language speakers describe the same image. If the model learns speech signals from language A that correspond to objects in the image, and learns the signals in language B that correspond to those same objects, it could assume those two signals -- and matching words -- are translations of one another.
"There's potential there for a Babel Fish-type of mechanism," Harwath says, referring to the fictitious living earpiece in the "Hitchhiker's Guide to the Galaxy" novels that translates different languages to the wearer.
The CSAIL co-authors are: graduate student Adria Recasens; visiting student Didac Suris; former researcher Galen Chuang; Antonio Torralba, a professor of electrical engineering and computer science who also heads the MIT-IBM Watson AI Lab; and Senior Research Scientist James Glass, who leads the Spoken Language Systems Group at CSAIL.
Audio-visual associations
This work expands on an earlier model developed by Harwath, Glass, and Torralba that correlates speech with groups of thematically related images. In the earlier research, they put images of scenes from a classification database on the crowdsourcing Mechanical Turk platform. They then had people describe the images as if they were narrating to a child, for about 10 seconds. They compiled more than 200,000 pairs of images and audio captions, in hundreds of different categories, such as beaches, shopping malls, city streets, and bedrooms.
They then designed a model consisting of two separate convolutional neural networks (CNNs). One processes images, and one processes spectrograms, a visual representation of audio signals as they vary over time. The highest layer of the model computes outputs of the two networks and maps the speech patterns with image data.
The researchers would, for instance, feed the model caption A and image A, which is correct. Then, they would feed it a random caption B with image A, which is an incorrect pairing. After comparing thousands of wrong captions with image A, the model learns the speech signals corresponding with image A, and associates those signals with words in the captions. As described in a 2016 study, the model learned, for instance, to pick out the signal corresponding to the word "water," and to retrieve images with bodies of water.
"But it didn't provide a way to say, ‘This is exact point in time that somebody said a specific word that refers to that specific patch of pixels,'" Harwath says.
Making a matchmap
In the new paper, the researchers modified the model to associate specific words with specific patches of pixels. The researchers trained the model on the same database, but with a new total of 400,000 image-captions pairs. They held out 1,000 random pairs for testing.
In training, the model is similarly given correct and incorrect images and captions. But this time, the image-analyzing CNN divides the image into a grid of cells consisting of patches of pixels. The audio-analyzing CNN divides the spectrogram into segments of, say, one second to capture a word or two.
With the correct image and caption pair, the model matches the first cell of the grid to the first segment of audio, then matches that same cell with the second segment of audio, and so on, all the way through each grid cell and across all time segments. For each cell and audio segment, it provides a similarity score, depending on how closely the signal corresponds to the object.
The challenge is that, during training, the model doesn't have access to any true alignment information between the speech and the image. "The biggest contribution of the paper," Harwath says, "is demonstrating that these cross-modal [audio and visual] alignments can be inferred automatically by simply teaching the network which images and captions belong together and which pairs don't."
The authors dub this automatic-learning association between a spoken caption's waveform with the image pixels a "matchmap." After training on thousands of image-caption pairs, the network narrows down those alignments to specific words representing specific objects in that matchmap.
"It's kind of like the Big Bang, where matter was really dispersed, but then coalesced into planets and stars," Harwath says. "Predictions start dispersed everywhere but, as you go through training, they converge into an alignment that represents meaningful semantic groundings between spoken words and visual objects."
"It is exciting to see that neural methods are now also able to associate image elements with audio segments, without requiring text as an intermediary," says Florian Metze, an associate research professor at the Language Technologies Institute at Carnegie Mellon University. "This is not human-like learning; it's based entirely on correlations, without any feedback, but it might help us understand how shared representations might be formed from audio and visual cues. ... Machine [language] translation is an application, but it could also be used in documentation of endangered languages (if the data requirements can be brought down). One could also think about speech recognition for non-mainstream use cases, such as people with disabilities and children."
Published September 2018