Friday, May 15, 2015

Philosophia


I just hit a new rock: Philosophia.  54 miles in diameter, 150 million miles away right now.  Apparent magnitude 12.7.

I have 3 exposures of 10 minutes each, taken with  t27 .  I zoomed way in on the center of the images, and did my slice-to-8-bit trick so that you are only seeing the 256 grayvalues that are just above the background mean.

It looks like I'm going to need much dimmer targets.



So that's nice -- the more images of real rocks the better.
But!  I may have hit the jackpot here.

What the heck are these things?



They look like rocks! 

Did I get a couple faint ones by chance in the same field of view?  It sure looks like it!   Maybe I can identify these gadgets!

If so, I can find what their apparent magnitude was, and that will provide a very valuable data-point for seeing how dim we ought to be able to go.


Thursday, May 14, 2015

Breakthrough

No, my next step is not to make a freaking testing system.

What I want here is an image processing and machine vision solution to finding faint streaks!  I don't want a testing system!  (Well, OK.  A testing system will be useful.  But first I want a solution to test.)

For months now I have been bouncing back and forth between a vision-based approach and a purely statistical approach.  At various times I could convince myself that one or the other was the Right Stuff -- but that conviction never lasted longer than a day or two.

The uncertainty has been painful.

At last I have a solution that has blended these two approaches: vision/structural, and statistical.  And I think this one is truly The Right Stuff.


Outline of the Approach

  • Get statistics for the background
  • Make the stars go away
  • do region-growing on all remaining pixels brighter than 2 standard deviations above background.  (experiment with this threshold.)
  • take statistics of the region size
  • see if there are any regions that greatly stand out

The region-growing is the machine vision part.  Using the statistics of the region size is .. um ... the statistical part.  My first experiment shows that these two together might be a Really Big Deal.

Region Growing

Does everybody know what region growing is?  It's easy.
  • You have a binary image, and you want to grow regions for the white pixels.
  • search the image in scan-line order until you find a white pixel.  Start a new region data structure and put this pixel in it. 
  • look at all its nearest  neighbors.  If you don't find any white pixels, you're done.
  • if you do find white pixels, add them to the region.
  • now check all the neighbors of the points you just added.
  • keep going this way until you run out of newly-added points to check around.  When that happens, your region has finished growing.


Will This Work?

I decided to check the last part first, because I already know I can make the stars go away.  But will the vision-and-stats part work?  If not, let's stop right here.
Let's go through the steps.





1. Simulate the Background

This is easy.  I have already measured the background statistics in several images and written a little gadget to make images with identical background stats.
Here is what one looks like, zoomed in on the relevant gray values.  (And zoomed in spatially to show nice big pixels.)



These 16-bit images actually look perfectly black.  What I have done here is to take a 'slice' of 256 gray values and display them in an 8-bit image.  The gray values are selected so that the darkest pixels in this image are about 4 standard deviations below the mean, while the brightest are about 4 standard deviations above.


2. Add a Streak

Next we add a simulated streak to the image.

The amount of energy that this streak adds to the image is determined by my studies of stars of known brightness in real images I have taken.  So this streak is  a simulated object of magnitude 20, and I am moving it 10 pixels.

If this were one of my real 10-minute images from T27 (  http://www.itelescope.net/telescope-t27/  )  that would mean an angular motion of about 0.5 arcseconds per minute.





The streak is right in the middle of the image, and slopes up to the left at about a 45 degree angle.

This is not what you would normally call a bright streak.   I think it would be very hard to find, by any normal means, in a 3056x3056 image, like the ones I get from T27.

It is miserable.  Hopeless.  Inconceivable.  We should give up.



3. Grow the Regions

These 'regions' are also called connected components, by the way.
Threshold at 2 standard deviations above the background mean (I should experiment with that) and find all connected regions of such pixels.

For debugging, I also draw all the regions I find into a new image, making each region white on black, for visibility.

Here is what we get:





4. Use the Statistic on Region Size


So here's the cool part.  Regions that are caused by random agglomeration of bright pixels certainly do happen, but the size of such regions has a pretty good standard deviation.

Their average size is about 12, with a standard deviation of only 2.6.  This means that it is very hard for such random agglomerations to get very large.  Only about 1 in a thousand of them will be larger than 20 pixels!

But the asteroid streak, even though in gray scale it looks awfully dim -- can grow as long as it wants!   So in this domain -- the domain of the size of regions significantly brighter than the mean background -- this thing is ... really big.



5. Be Shocked and Awed


In this domain, the size of the asteroid streak's region is fifteen standard deviations above the mean.

In technical statistical terminology, that is Freaking Enormous.

If we were talking about random variations in audible noise, and a fifteen standard deviation increase occurred -- it would blow me across the room and through the wall.

I think we may be onto something.



Friday, February 6, 2015

Image Processing, Machine Vision, and the Way Forward.




There are no stars in this image.




There are also no galaxies, no asteroids, and there is no interstellar background.



Images contain  pixels, and nothing else.  The image can answer questions like "what is the gray value at location x,y?".  Or:  "What are the dimensions of this image?"   Or "what is the average gray value?"  Or even "What is the gray value histogram?"  That is all.

If something says "I see stars in that image."  then that something is a vision system.  The stars that it sees are not in the image -- but are in its own head.  In the data structures that it has created, by doing some highly nontrivial computation, using that image as input.

It's very hard for humans to understand that the objects that they see are not explicitly in the image, but are highly abstract constructs in their own minds -- because, for humans, vision is utterly effortless.  When you look at this page and then glance around the room, you are probably using more compute power than exists in the United State of America, including all the secret parts, and you're doing it without so much as frowning.  (Which makes it pretty hard on guys who try to get machines to see what you can see.  But that .. is another story.)


Image Processing is not Vision


There are two kinds of processing you can do to an image: image processing and machine vision.

Image processing consists of operations that take an image in, and put out a transformation of the image.  For example, an image processing operation may take in a color image and put out a monochrome version of it, or a contrast-enhanced version.  Or it may put out the sum of all the pixel values, or the average pixel value, or a histogram of all the pixel values.

Image processing stays within the realm of images and their properties.

Vision, on the other hand, takes in an image, but then puts out a data structure that is not in the realm of images-and-their-properties.  For instance, it may take in a gray scale image and put out a data structure that says "I see a chair at position X1, Y1, X2, Y2 with certainty C".

Chairs are not in images.  The vision system needs to add a lot of non-image knowledge to determine that a certain pattern of brighter and darker regions probably represents a chair.

But even simple features like bright edges, or streaks, or regions of a given color are not quite in the image domain.  A uniformly colored region is not explicit in the image.  It is implicit, and must be gotten out  (turned into a data structure)  by some nontrivial processing.

 

I Need Both

In the asteroid-finding system I am writing, I should have both an image processing level and a machine vision level.  These two levels should be well separated from each other, so that I can run the lower-level image processing by itself.

Because -- I want to be able to test the machine vision level against what a human can do, after only the image processing code has been run.  The goal will be to get machine vision code that can at least get in the ballpark of a system composed of image processing plus human vision.



Testing


To characterize vision-system performance I will make a little gadget that creates random simulated asteroid streaks in real images that I have taken.  

  • Take 50 real images.
  • program draws random streaks into half of them.
  • Starting point and direction are random.
  • brightness of streak is not random -- chosen by input argument.
  • I don't know which images have streaks, which don't.
  • I examine all 50 images visually, write down locations of where I think I see a streak.
  • Look at 'answers' saved by streak-drawing program, count how many real streaks I missed, and how many times I thought I saw a streak that was not really there.  ( False negatives and false positives. )

I should probably have some description of what kind of performance I hope to achieve with any vision system, whether it is natural or artificial.  I.e. what level of false positives are acceptable: when I hallucinate a streak. 

The goodness of the vision system should probably be expressed something like this:  "For every four streaks that the vision system reports, three of them, on average, will be real."

In this endeavor, false positives are very expensive.  They cause you to go take a follow-up picture.  False negatives are not a big deal: when you miss a streak that is there.  After all, the point is finding new asteroids.  If we miss one, that's OK.  We didn't know about it before, and we still don't.

Well, unless I miss the one that's coming to destroy civilization, or our species, or life on land, or whatever.  That would be bad.


Next Step


OK, so that's my next step.  Write this testing system, and use it to characterize my own 'natural' vision system performance.  See if it also gives me ideas about improvements to the image processing level.  And get ready to use the same test on the machine vision system.







Tuesday, December 30, 2014

Tantalus 2


So I'm claiming that sliced images is the best thing since sliced bread.  Well if I recall correctly, the image I used as an example in yesterday's post actually has an asteroid streak in it.  (In fact, it has the only decent streak I have yet acquired, but we'll get to that later.)

So -- what effect did this slicing process have on that streak.

Let's zoom in.



The streak is there, but the central part of it is too bright for this image, and has been blacked out.

That's good news!  The brightness of Tantalus when I took that image -- the 'apparent magnitude' -- was 17.62.  I want to be able to see streaks that are much dimmer than that, and this result suggests that dimmer streaks might work very well.

And there's a way we can estimate how dim we might have been able to go with this telescope, on that night.  We can simulate our own streak, in this very image.


The Streak Simulator

I haven't told you about how I do this software yet, and -- I probably won't.  It's not terrifically interesting, compared to pictures of rocks in the sky. 

OK, I'll make it quick.  I write my image processing software in C, on a Fedora 20 system with a GCC compiler.  I write it all from the ground up, using no ancillary image processing libraries.  I like it that way.  Programming the bare metal.

I use the tool-building philosophy of all intelligent programmers: make simple tools that do a single job well, and combine them together with a scripting language.  Bash, actually.  (Until recently I used csh, which means that I got started doing this stuff when the universe was quite a bit less red-shifted than it is now.)

So -- the streak simulator is a program I wrote recently.  You tell it how much total energy you want it to deposit on the image, where the x,y start point is, what the direction of travel is, how far it should go, and how many seconds that should take. 

The program then moves its idea of where the asteroid is in tenth-second increments, at every moment doling out its increment of energy in a randomly-chosen direction, and at a random (normally-distributed) distance from the asteroid's 'true' position.

It's not perfect, but it's pretty close.  Here is what it did when I used it to try emulating the real streak that Tantalus made in my image.  (The simulated streak is just to the right of the real streak, and I made the simulated one perfectly vertical.




That looks pretty good!  Except I had to put 45,000 grayvalues of brightness into it, when the real streak only used 31,000.  Hmm.   And it's still a little scrawnier-looking than the real one.   Hmm.  

So, I don't know how good a model this really is, but just in case it is predictive, here's what it predicts:



If that is really what a mag 19 streak looks like, I think I can detect that like falling off a log.  And this is with a 20" telescope!  Through iTelescope.org, I have access to a 27" that collects twice as much light.  With that, I might hope for mag 20, or better!

But!  Simulation is one thing.  Ground-truthed (so to speak) data is another.

What I really need now is more images of known rocks, at known brightnesses.


Monday, December 29, 2014

Out of the Wilderness

I've wandered far since my last post, lost in the dark places of the southern sky and lost in the myriad possibilities of the things that can be done, trying to distinguish them from the things that should be done.

I think I understand better now, and hopefully even well enough to explain.



A hunter does not simply load his rifle and go running into the brush waving it about.  A hunter makes a plan before even choosing the weapon, and the first part of the plan must be: What am I looking for?


Well, what I am looking for is very faint streaks -- streaks that are just barely above the background noise.

OK, good.  Now we're getting somewhere.  This clear statement of purpose immediately suggests a question.  What is the background noise?  What does it mean to be 'just barely above it'?


To answer that, let's look at my raw image again.





Let's see.  How will we determine what the background of this image is?  Using my years of training in machine vision, together with an innate talent for noticing facts that are glaringly obvious, I soon determine -- that this image is all background.

All we need to do is look at a histogram of the image, select the most popular value, and we will have found the mean of the background.

So, let's get the histogram.





The red parts are the plotted points, showing how many pixels were found in the image at a given gray value.  That spike to the far left is so large that it makes the rest of the plot look flat.  That's what I mean by an image that is "all background".


Let's zoom in on the part of the histogram that only shows the darker values that are in the background.

Here is a graph of just the darkest 4% or so of possible grayvalues.





Now you can see some detail.  The background, that simply looks black in the image, actually makes what looks like a nice, normal distribution whose mean is just below grayvalue 1000.  Practically all of the pixels in the image are represented by the spike we are seeing here.  This is the distribution of the background.

Let's look a little closer yet.




That's about as nice of a normal distribution as you will ever see.  Its mean looks like it's at about 940, and the point at which it falls to half that value looks to be about 900 on the left and 1000 on the right, which means that the "Full Width at Half Maximum" of this distribution is 100, which means that the standard deviation is about 42.    (The FWHM of a normal curve is always 2.35 sigma.)

But what does this tell us about how we can view the pixels we care about?



What it says is that we can do something beautifully simple.  Make an eight-bit image with its paltry 256 gray values simply centered on the highest point of that histogram.  That will show us pixels that are 3 standard deviations above and below the mean, which should be plenty.

I call this a 'slice' image, because I am slicing out 256 gray values from the 65,536 possible from the 16-bit original.

Here's how we do it.  Look at the peak of that histogram at 940.  Go 128 below that, to grayvalue 812, and make that our zero point.

Now make an 8-bit image as large as the original.  Go through and subtract 812 from every value in the original.  Any pixel that would be below zero, we just set to 0 in our new image.  And any pixel that would be above 255, we also set to 0.

Because we don't care about those pixels.  They are not close enough to the background to be part of the really faint streaks we want to find.

So what does the result look like?
Let's see.




We just got rid of all the stars.  They have turned black because they are too bright to be in our new 8-bit image.  What you are seeing now is a really nice smooth picture of just the background, against which we will be able much more easily to find our faint streaks.

SInce May, I have been wondering how in the heck I could make the stars go away, and here it is in a simple thresholding operatrion that took all of 134 milliseconds of a simple processor on my laptop -- and that was for the full sized original image that is 3056x3056.  That is a dirt-cheap operation.

This job just got a lot easier.



Sunday, May 25, 2014

Tantalus, my First Rock





My first rock, 1.5 miles across, 85 million miles away, images from the iTelescope T24 in Auberry, California about 2 hours ago.

The image is processed down to 8 bit gray with my own software -- in the original you cannot see the streak.

These images are zoomed way in.  They are both about 250x250, while the originals are 3056x3056.


At 0200 Auberry time, 0500 Michigan time, 0900 GMT -- a ten minute exposure:




And at 0211 Auberry time, 0511 Michigan time, 0911 GMT:




Si muove!


Oh, we are going to have a lot of fun with these two images.

By the way, here is what the images looked like before being 8-bitified so you could see faint gray values better.  You can see the brightest few stars, and that's about it.




And, say hello to my little friend, who took the picture for me, made by PlaneWave Instruments, owned, adapted for the internet, and very well managed by iTelescope.net, the lovely and talented T24, of Auberry, California:








Tuesday, May 20, 2014

Into the Dark


I propose to use none of the usual methods to find rocks.  I propose to use a method that requires such ridiculously vast amounts of compute power that it's never been tried before -- because the compute power just wasn't available.

We will be looking in our images for streaks so faint that no amount of contrast enhancement can make them visible.  So faint that they can only be detected by some combination of statistics and machine vision.

But to do that -- we will need to understand the dark between the stars.



So first, let's go find some.

Here is the image I showed a while ago, that contains the Hubble Deep Field in its center.




OK, so I don't want to use the center, because I actually was able to see some galaxies in there.  I want to find a rectangle that has no discernible stars even after serious contrast enhancement.  And I want the rectangle to be as big as practical, because we are going to do statistics on its pixels.

After some searching around, I find this nice little spot right here:




Let's look a little closer.




Okay!  No stars that I can see, after doing the best contrast enhancement I could get out of the Gimp.

( Note!  In all of the work I have done or will do on this blog, I use only free public websites, free open source software, or software that I have written myself. )


So, we found a dark place.  What do we want to know about it?

What I hope to see is that those dark pixels, statistically, have a nice normal distribution, and spatially, are nice and smooth, with no discernible clumps.

It's easy to get the statistics from that region.  Here they are:
  
          count: 71416   
     min:       0.000 
     max:    1513.000 
     mean:    972.074 
     sigma:    37.493

The images I get from these telescopes are 16-bit grayscale, so the total range of pixel values is from 0 to 65535.   So an average brightness of 972 is pretty darn dark, which is why that rectangle looks plain black in the original image.



Okay, the stats will be useful, but they  don't tell the whole story.  What do those pixels actually look like?  I can't see them now because display technology can't show 16 bits of grayscale, and human eyes wouldn't be able to see it if they did.

So here's what we can do.  The pixels in that rectangle go from 0 to 1513.  So divide them all by 5.9, and that will make them all be in the range 0..255, and then we will able to see them properly in a normal image!

Here's what that rectangle looks like after this treatment:



A few little smudges in there, but basically nice and smooth!

And finally, I would like to know if those pixels are more or less normally distributed.  Because a lot of the reasoning I will do will depend on that.

After going through all kinds of pain to try and find a test I could use to determine whether or not a bunch of numbers are indeed normally distributed, I have finally settled on the simple expedient of graphing the darn things and looking at the graph.  Just print out the values, pipe them through "sort | uniq -c", and use a nice little gadget called gnuplot.

Here's the result:




That's glorious.  That curve looks like a textbook illustration of a normal curve.  On the Mick's Arbitrary Graphical Expedient Test of Normalcy, that curve gets a score of 0.995, which means "Heck Yes That Is Normal!"


This all means that the Dark Between the Stars is just how I hoped it would be.  It will make a perfect hunting ground in which to seek the faint tracks of our prey.