Saturday, September 21, 2019

Things to Think About When Deciding Whether Or Not to Vote

Election day is fast approaching. With the increase in polarization between our primary political parties, many people have adopted a dim view of participating in the process. The importance of voting seems to be lost upon those who view the entire political process with contempt. If I may, following are some views I have come across to which I would offer an alternate viewpoint.

"I just don't feel like there is a candidate that represents my beliefs."

I say this with real respect and understanding for your position, it's not about you. We don't (or at least should not) vote in order to get our positions at the top of the list. It's about all of us. Those who we elect are going to make decisions that affect the homeless, veterans, cancer patients, the wealthy, blue collar workers, artists, doctors, and the list just goes on. Many of these groups may not touch your life directly, but they impact the communities we live in, our culture, and our shared values.

"I don't feel like my vote counts."

It is true that you are just one of millions when you vote and that can certainly make one feel insignificant. However, if you don't vote, it counts even less. If you do vote, you join with your neighbors, coworkers, family, friends, and the rest of the country in setting the direction, even if only a small amount. In addition, during my time spent in county government, I participated in the vote counting process. I've seen a number of elections where the margin was just a few votes. You may end up being one of just a few that launch a future, well-known leader into a political career.

"It's all just a corrupt bunch of people in Washington or (insert your state capital here)."

To be sure there is a lot of corruption, and if not corruption, certainly a lot of self-serving behavior. Let's have a little thought experiment: Jim believes that politicians are corrupt and he chooses not to vote because of this. Jim's viewpoint implies that he believes his behavior is more acceptable than that of the corrupt politicians. It would be a conflict for Jim to see himself as more corrupt yet at the same time apply that label to others. Now if all the people like Jim refuse to vote for the same reason, the people who do vote, don't oppose the corruption? When we don't vote because we don't want to support a system filled with corruption, we end up promoting that same corruption.

"There's plenty of time. There's an election at least every two years."

In fact, there is not plenty of time. One of the interesting side effects of living in large communities is that change tends to be slow. We often want this thing or that thing in society to change right away. However, slow change is useful in that it gives society time to make incremental change, evaluate the results, and do a course correction before we go too far. The decisions made by politicians today will have effects that will last for years or may take years to realize. Consider the potential longevity of members of Congress or the Supreme Court. Individuals in these bodies are often in their positions for years. The decisions made elected officials today will affect your future and your children even if you don't have any yet.

"I don't have time. It's too inconvenient."

I completely agree! It is too inconvenient. There are initiatives to make voting easier; things like making election day a national holiday and moving it to warmer months. But for these initiatives to take hold, decisions have to be made and those decisions are made by elected officials; the people we vote for. If you don't vote, if you don't take this concern to our policy leaders, it will never change.

"Elections are rigged."

We've heard a lot about this lately. Now that some of the dust has settled on the 2016 election where this claim was thrown around so often, take a moment to look back at the results of the 'rigged election' claims. How many have actually resulted in criminal investigations that have found wrong doing? A Heritage Foundation study tracking voter fraud identifies 1088 instances with 949 criminal convictions. That seems like a lot, but reading the report gives some interesting insights. The vast majority are for local (City/County/Township) elections. Almost all impact a small number of ballots (relative to the number of total ballots cast). In many of the cases, the proven instance is a single person who committed fraud for only their own vote. Generally, the margin that a race is won by is substantially larger than the number of questioned ballots, meaning that the fraud had no impact on the outcome of the race. Finally, in every single allegation of 'millions of fraudulent ballots' being cast, no evidence was produced to support the claim. Many people got their names in the headlines, but no reliable evidence has turned up.

Wrapping it Up

Having the ability to vote carries with it much more impact than just trying to get one's candidate in office. It is one of the few ways that most Americans can directly participate in the how this country is run. When we all vote, it sends a message to those in office that we are interested and we are watching. We are taking action against those that fail to live up to their responsibility as an elected leader. It can benefit you, but it does benefit all of us. Please consider taking a few minutes to learn just a little about those running, head out to the polls, and cast a ballot for the one you think is the better candidate. We may not get the best one for the job every time, but if we keep electing the better candidate year after year, we will get there, and many of these concerns above will start to fade away.

Sunday, December 2, 2018

Day Hike: Red Rock Canyon Conservation Area

Earlier in the year, I was offered the opportunity to attend a technical conference in Las Vegas. After a bit of scouting around the area, I came across Red Rock Canyon Conservation Area, and found some great hiking trails. Following are the trails I picked for the day.

The day started by waking up in Sandy Valley in one of the many free camping areas set up along Sandy Valley Road. The hope was to get some nice images of the Milky Way, however the full moon ended that. I did learn that there is a name for this free camping: boondocking. Essentially, this is the practice of camping for free on public lands. It's not unlike camping in a campground, but without the amenities. It's perfectly legal on established sites, but the same 'leave no trace' rules apply. It was incredibly cold, but still a beautiful sunrise.

The first light over the mountain range just south of Mineral Springs. Sirius appears as the brightest object in the sky.
Mountains north of Hwy 160 in Sandy Valley.

I then drove to the conservation area, and paid $15 for my day pass. I parked in one of the parking lots, and readied my gear. I hit Whole Foods the night before, and bought about a gallon and a half of water to fill my bottles and camel bag. It was then time to start. As a side note, there are clean water fill-up stations at the visitor center, so bringing your own water is not a necessity.

There is a trail head on the west side of the picnic tables (situated next to the upper parking lot) that provides a nice start. It takes you west, then curves around north of the visitor's center, then crosses the road toward the three Calico stops. Hiking next to this ridge was really interesting. There are multiple laysers of rock here, some red-orange, some light tan, and some gray. There is a bit of up and down along this trail, but nothing more than 10 - 20 meters. There is a small bit of scrambling, but I can't see calling this section difficult. One tip: when the path ends up going over the rock, look for a washed-out path along the rock. This is where dust from the rest of the trail is deposited, and it acts as a guide when walking over the rock.


Starting the hike, heading toward the first Calico lookout point.

Each of the Calico stops (Calico 1 and Calico 2) are beautiful. There is a parking lot for each, however no infrastructure. The longer section after Calico 2 takes you to the third stop, which includes restrooms. Heading north from that parking lot, you are quickly met with a sign: right for the Calico Tanks, and left for Turtlehead Peak. I started out by heading toward the tanks.

This was essentially walking through a ravine between two ridges to the north and south. As you get deeper in, you end up scrambling to multiple beautiful lookout points. From there, you can see Las Vegas, as well as the visitor's center where we started out. Looking behind, you can also see Turtlehead Peak. The return from the tanks is essentially retracing our steps out. It's important to be careful here, as sand is deposited all over the rock by the multitudes of tennis shoes going in and out of the tanks. That makes for lots of opportunities to lose your footing.

Hiking up in the Calico Tanks with ridges to the left and right.



A view of the visitor's center wehere I started out that morning from one of the lookouts on Calico Tanks.

Back at the sign, I turned right, and headed up toward Turtlehead Peak. At first, the climb wasn't too bad. There was a slight increase in elevation as we walk through a washout to get to the base of the mountain. Then the increase becomes more pronounced as we start to ascend up the sand and gravel mixture that has worn off the mountain over the past few hundred thousand years. At some point, we start seeing sharp, jagged rocks protruding from the ground, and the scrambling begins. This stretch was for me, the hardest part of the climb. This will get you up to the ridge very close to the peak. The entire trail is marked with orange paint (usually either a circle to indicate that you are on the path, or an arrow telling you which way to turn). Once we hit the ridge, the incline is not as bad. You work your way around to the north side of the mountain, and there are trails that follow around the peak with a slow incline, or you can head straight to the top. I selected the longer but less inclined route.

A view north from the ridge of Turtlehead mountain.


Looking at Google Earth, and other maps, it appears as if the trail stops a few hundred feet short of the peak. This was not the case. If you keep an eye out, you will see small straight lines or two dots side by side in a bluish-grey color from a spray paint can. These will guide you to the top, though at this point, it's not really necessary, as it's obvious how to get to the peak.

A view of Las Vegas, the visitor's center, and the Calico Tanks from Turtlehead Peak.

At the top, once I slowed down, I noticed it was significantly colder. It was probably in the upper 20s or low 30s at around 2:30 p.m. Again, the views were stunning. The multicolored rock really popped out, and the surrounding mountains quite a bit to take in. At this point, I have hiked roughly 8 miles with over 3000 feet of cumulative climb. I was extremely tired. Unlike the Bald Mountain climb, there was plenty of oxygen (altitude was only 6323 feet). That said, the steep stage of the climb really hit me hard. After a few pictures, I hooked up with a group of friends from the area doing their 'regular weekend hike' for the hike down.

Made it!


Here is an important point about hiking up mountains. Hiking down is definitely easier, but it is in no way, easy. I started to get pretty tired, especially on the steep parts as it took a lot to keep myself balanced, and stop myself on each step down. By the end of the hike, my left knee was pretty sore, and remained like that for two days. This is a great place for some trekking poles to take some of the shock in stepping down.

When I got back to the parking lot, I chose to hike along the road back down to the visitor center. I was hoping to hit Icebox Canyon, but time did not allow. I took the road all the way back to the visitor's center. The hike down was 1 1/2 hours.

There is no way to describe the beauty of this area. The mountains are larger than anything I've ever seen, and stand silently, daring you to try to climb them. The sky is a beautiful shade of blue, unmarred from air pollution. The air was clean, and the silence was amazing. This is easily one of my favorite hikes ever, and definitely worth repeating, should I ever be back in the area. Here's a short video of the hike:


The red marks the trail I followed. It starts in the lower-right corner.






















Elevation profile for this hike. The first peak is Calico Tanks, the second, taller is Turtlehead Peak.


Details:
Time: ~7 hours
Distance: 12.24 miles
Cumulative Climb: 3492 ft.
Min Elevation 3704 ft.
Max Elevation: 6323 ft.
Temp: 20s - 42 deg. F

Friday, August 31, 2018

Hiking the Dunes: Warren Dunes State Park

Finally, after a waiting over a year, I've made it to Warren Dunes State Park. The drive was longer than I'm used to, but it was worth it. First tip: Bring cash to the main gate. The alternative is to be directed to the campground entrance where cards are accepted.

Once in (and having paid at the campground office), I headed back toward the main gate and turned off early into a parking lot. This is the trail head. I started out entering the trail, and taking a right (to do the loop counter-clockwise). The wooded area was humid, and very flat. Insects were not really bothersome, in spite of all the recent rains. Speaking of... There were a number of muddy areas along the trail, but most had a small path going around.



I decided to take trail 11, which should have looped back to the same spot, at which I would head north. Roughly 1/3 mile in, I realized that the trail was not curving back around as it did in the map. At that time, a hiker came by and told me of a really interesting undocumented trail that went over a hill to a water tower, and finally ended on the beach. I decided to try it. I made it over the hill (good workout, by the way), but the entire trail was covered with spider webs. Every 40 or 50 feet, I would walk into another. When the trail started to become obscured by overgrowth, I elected to head back the way I came, and just turn right to head toward the beach.


After a bit more woodland, the trees thinned out into a mostly sand covered area. From that into sand dunes spotted with maram grass. I found my way to a tree in the valley between two dunes, and sat for some shade, rest, and water. After a protein bar and some water, it was back to it. I hit the lake shore, and the temperature cooled a bit. There was also a nice breeze. When I walked inland back to the trail, the temperature seemed to climb 15 degrees (which I withstood for another thirty seconds until I could get my self back to the shoreline).


Once I hit the parking lot, I turned south to head back. I immediately started up one incline, then another, and another. I ended up climbing the edge of the dune on the north end of the parking lot. After a walk through a small barren area (only sand), I headed down to find a shady spot again to rest. At this point, I was starting to suffer from symptoms of heat exhaustion. Despite the fact that I wanted only to get back to the truck into some air conditioning, I stayed there cooling off until I felt rested.

The next part was completely off trail. I stayed in the valley between the two dunes so I could minimize climbing. I also headed straight for the nearest set of trees that were back toward the trail head. I ended up topping one more dune, crossing the a stream, and doubling back about a half mile, but I was back at the trail head, dumping the sand from my shoes.

And in all of that, nature finds it's slow but certain way, quietly. Each step shakes lose a small bit of the worry of the day. And I am in awe of our tiny existence in the universe.









DETAILS
Time: 4 hours
Distance: 4.98 miles
Cumulative Climb: 682 ft.
Min Elevation 560 ft.
Max Elevation: 754 ft.
Temp: 95 deg. F

Sunday, August 5, 2018

Day Hike: Devil's Lake State Park

I took a day to head out to Devil's Lake State Park in Wisconsin to do some hiking. This one was the longest, and most difficult of all of my day hikes, thus far. It had some of the most beautiful views, but also some points that were a bit discouraging.

After paying for parking, and getting a trail map at the visitor center, I started off about 7:30 a.m. I picked up the West Bluff trail head in the main park area. From here it was straight up for about 480 feet. The remainder of the West bluff was up and down for a while along the west side of the lake. Views were really beautiful, and the trail was entirely paved (or natural rock). A steep decline put me on the south edge of the lake. There was a good bit of walking next to a road until I reached the southern park area. That took me to Grotto's Trail. This was a nice area, walking along the base of the southern face of East Bluff for a while. The entire time you would hear traffic from the same road I walked along earlier. Then it was across some railroad tracks, and a steep incline up.

On West Bluff looking southeast.

The hike up the southern face was very hard. It was mitigated with some stairs made from the quartzite rocks, but still a very steep incline. I passed some other hikers, and sport climbers along the way. A portion of this climb is in the video:





Once at the top of the east bluff, I took a break and prepared for the hike east on the Ice Age trail. Unfortunately, the map provided in the visitor center incorrectly had me go east right above the incline. This started taking me along what appeared to be a trail, but eventually turned into nothing more than a gulley carved out by rainfall runoff. I ended up doubling back, and up around 150ft back to the top of East Bluff. More wandering around, and I eventually found the Ice Age Trail by way of the East Bluff Woods Trail. This entire area is littered with signage, but there are a number of unmarked intersections which leave you searching for which trail you want to be on. This area is also full of other hikers and tourists there for the views of the lake.

Atop the southern face of East Bluff looking east-southeast.


Once back on the Ice Age Trail, I headed east, and back down the mountain to Roznos Meadows. This was a nice leisurely stroll with a welcomed flat surface. A few small inclines, and nice views of the southern face of East Bluff puts you at a trail head along highway 113. It was at this point where I realized that I would not have time to make it to Parfrey's Glen (the original eastern-most destination), so I took 113 North to another Ice Age trail head. This was, in a word, grueling. It was over a mile up from 860 to over 1300 feet. In addition, walking along the highway was hot with little by way of a comfortable place to rest. I finally hit the trail head, and headed back into the woodland.

Looking at the south face of East Bluff from Roznos Meadows.


My first stop was just inside the wooded area where I stopped at a huge log for some lunch and a rest. The break and refueling did a lot to re-energize. I recall distinctly the apple which was not only sweet, but very cool. Notwithstanding some ups and downs in elevation, this was generally a long downhill portion of the hike, taking me down to the campgrounds back in the park's main area. It also took me to another poorly marked part of the trail. I ended up working my way through the entire campground, missing where the trail headed back to the parking lot where I parked.

All in all, I would probably like to try the trail again, but with less of a time constraint (I had to be back at the hotel by 3:30 p.m.) I would be able to skip the highway 113 portion, and hit Parfrey's Glen. The views were beautiful, and once you picked up the trail east from the south face, there were no people at all.

The 15.3 mile loop. Darker color is higher elevation, gree (in Roznos Meadows) is the lowest elevation.

The elevation profile. The first incline was up West Bluff. The second is Grotto's Trail, then up the south face. The find steep incline is up highway 113.

Details

Min. Altitude: 797 ft.
Max Altitude: 1528 ft.
Cumulative ascension: 3024 ft.
Distance: 15.3 miles
Duration: 6 h 50 min.
Temp: mid 70s - upper 80s F.

Saturday, June 2, 2018

My First Book

Well, this was a long time coming (since 2007, actually), but I finally finished my first book. Really more of a story, it's only 41 pages, and that includes appendices. But it's finished.



I started this in 2007, in a search to find out, as best I could, what happened to the B-17 carrying my great uncle Richard E. Hargrove in their flight from Gander, Newfoundland to Warton, England in 1943. They were lost early in the flight with no trace.

After meeting family of other crew members, and a navigator that was part of the group flying over that night, I pieced together, and analyzed the evidence, which resulted in this story.

One of the nice things about researching a topic like this is that going in, I felt I knew what happened, but after a detailed review of reports, communications, and stories, I have completely changed my view of what I think is the most probable reason they went missing.

From the Contents:
Introduction
The Crew
Training
The Ferry
Aftermath
Epilogue
Appendix A: Report of Aircraft Accident
Appendix B: Missing Air Crew (MAC) Report
Appendix C: Next of Kin List
References

Thursday, November 16, 2017

Glenwood Dunes Trail System Hike

July 2018 Update: The quality of the video shot last November was so poor, I re-shot the hike on Memorial Day, 2018. The video below is from 2018.

Earlier this month, I filmed the Glenwood Dunes hike, a hike I first did last year. the video is finished, and now up on YouTube. Feel free to have a look:





This takes place in the fall, so the colors were out and very bright. I used a different camera for this hike, and I wasn't very impressed with the quality. I'll likely go back to my other camera for the next in the series.

Tuesday, November 14, 2017

Web Scraping 2: Some Intermediate Functionality

IMPROVEMENTS AND CHANGING BUSINESS REQUIREMENTS


In my previous post on web scraping, I described a very basic script to cull data from the Internet. That was a pretty good first attempt, and the simplicity of the web application, as well as the business requirements didn't require a great deal of coding. The follow-up however, required some additional work as the previous script would not handle many of the new issues that arose. In this script, we are working with data from the Apache County, Colorado Assessor's Office. While the web application is the same, there are some changes that require some additional work. Specifically, we address the following changes:

  • We need a better way to wait for elements to become available on the webpage. Adding time.sleep() statements is fine, but we can't always be sure that the number of seconds we specify will be sufficient, and the more time we specify, the longer the script will require to run, as time.sleep() is a simple wait command. We need something that waits for the page to load completely, and if that happens in .75 seconds, then processing should continue after .75 seconds. If it takes too long, a timeout should interrupt the script.
  • We need to add some logic to identify records that we do not want. For example, if there certain records that we are not interested in, we should break out of the processing for that record immediately, and conintue to the next. This reduces the file size, but more importantly, it reduces the run time of the script, and reduces the code path (thus eliminating the potential for errors to arise, and cause an unexpected failure).
  • We need a better XPath tool. After a recent update, the XPath generator tool we used in the last article stopped working.
  • We need a way to handle the possibility of multiple items in a list that might be returned from a single search.
  • The data we need might be split among multiple screens. We should be able to move between screens to capture all of the data we want.
  • Some of the data we need may be in a frame. We need a way to be able to select the frame which contains the data we are looking for.
  • We need a way to access objects which may not be visible on the screen.
  • Sometimes we hit a parcel ID that is not (no longer?) in the system. This produces an error that will stop the scraper and throw an exception. We need to gracefully handle 'parcel not found' errors.
  • Along with the previous item, it would be useful to log what takes place with each record from the original file. Adding a log file would allow us to record successes and failures, and the reason a given parcel could not be retrieved.

UPDATED CODE


The following code addresses each of the issues highlighted above. As with the previous post, we'll make notes in code, then explain in more detail below.

# 1. Imports
from selenium import webdriver
from selenium.webdriver.common.keys import Keys
from selenium.webdriver.common.by import By
from selenium.webdriver.support.ui import WebDriverWait
from selenium.webdriver.support import expected_conditions as EC

from selenium.common.exceptions import NoSuchElementException
import time
import re
import csv

# 2. Create a webdriver object.
driver = webdriver.Firefox()

# 3. Setup Input and Output files
# Input file format is: AcctNo,ParcelNo.
with open('allkeys.csv', 'rb') as f:
    reader = csv.reader(f)
    parcelList = list(reader)
totItems=len(parcelList) # get the count of total items for status.
# 4. Change the mode of the output file from write to append.
outfile = open('datafile.csv','a')
outfile.write('"Account No.","Parcel No.","Legal Class","Unit of Meas.","Parcel Sz.","Short Owner Name","Address 1","City","State","Zip Code"\n')

logfile = open('ApacheCoScraper.log','a')

# 5. A counter is used to tell us how far along we are. It's used below.
iCntr = 0

# 6. Read a row from the input file - this contains 2 fields now.
for row in parcelList:
    # 7. Load the County page and get past the splash page
    driver.get("http://www.co.apache.az.us/eagleassessor/")
    time.sleep(3)
    driver.switch_to_frame(driver.find_element_by_tag_name("iframe"))
    # 8. Scroll up to see access the submit button.
    submitElement = WebDriverWait(driver, 10).until(EC.presence_of_element_located((By.NAME, 'submit')))
    driver.execute_script("arguments[0].scrollIntoView(false);", submitElement)
    submitElement.click()
    # 9. Wait for the page to load (applies to remainder of script)
    driver.implicitly_wait(15) # seconds


    # 10. Search for a parcel number, and bring up the Account Summary page
    accountNumber,parcelNumber = row
    parcelbox = WebDriverWait(driver, 20).until(EC.presence_of_element_located((By.NAME, 'ParcelNumberID')))
    parcelbox.send_keys(parcelNumber)
    parcelbox.submit()




    # 11. Try to find a warning message (indicates 'parcel not found').
    try:
        warningMsg = driver.find_element_by_class_name('warning')
    except NoSuchElementException:
        pass
    else:
        iCntr += 1
        print "%s (%d/%d) not found." % (parcelNumber, iCntr, totItems)
        logfile.write(parcelNumber + '(' + str(iCntr) + '/' + str(totItems) + ') not found.\n')
        continue


    # 12. Multiple entries may appear, select the one we need.
    validRowLink = driver.find_element_by_link_text(accountNumber)
    try:
        validRowLink.click()
    except:
        driver.find_element_by_class_name("clickable").click()
   
    # 13. Assume we only want records with a type of "02.R (land only). Ignore anything that is not 02.R
    raw_Legal_Class = driver.find_element_by_xpath("//*[@id='middle']/table/tbody/tr[2]/td[3]/table[3]/tbody/tr[2]/td[1]")
    legalClass = raw_Legal_Class.text

    if not legalClass == '02.R':
        iCntr += 1
        print "%s (%d/%d) skipped." % (parcelNumber, iCntr, totItems)

        logfile.write(parcelNumber + '(' + str(iCntr) + '/' + str(totItems) + ') skipped.\n')
        continue

    # 14. Capture Account Summary page data (the first page of data)
    raw_Account_Number = driver.find_element_by_xpath("//*[@id='middle']/h1[1]")
    raw_Parcel_Number = driver.find_element_by_xpath("//*[@id='middle']/table/tbody/tr[2]/td[1]/table/tbody/tr[1]/td[1]")
    raw_Tax_Area = driver.find_element_by_xpath("//*[@id='middle']/table/tbody/tr[2]/td[1]/table/tbody/tr[2]/td[1]")
    # 15. Capture the data items as python varaibles before moving to the next page.
    accountNumber = re.sub('Account:', '', raw_Account_Number.text).strip()
    actualParcelNumber = re.sub('Parcel Number', '', raw_Parcel_Number.text).strip() # actualParcelNumber has '-' marks in it.
    taxArea = re.sub('Tax Area', '', raw_Tax_Area.text).strip()

    # 16. Jump to the Parcel Detail tab
    PDPageLink = driver.find_element_by_link_text('Parcel Detail')
    PDPageLink.click()
    # 17. Obtain page data as webdriver objects
    raw_Unit_of_Measure = driver.find_element_by_xpath("//*[@id='middle']/div/span[6]")
    raw_Parcel_Size = driver.find_element_by_xpath("//*[@id='middle']/div/span[8]")
    # 18. Extract data from the webdriver objects, and place in python variables.
    unitOfMeasure = raw_Unit_of_Measure.text.strip()
    parcelSize = raw_Parcel_Size.text.strip()

    # 19. Jump to Owner Information tab
    OIPageLink = driver.find_element_by_link_text('Owner Information')
    OIPageLink.click()
    # 20. Capture Owner Information page data as webdriver objects
    raw_Owner_Short_Name = driver.find_element_by_xpath("//*[@id='middle']/div/span[2]")
    raw_Address1 = driver.find_element_by_xpath("//*[@id='middle']/div/span[6]/table/tbody/tr[1]/td/span[2]")
    raw_City = driver.find_element_by_xpath("//*[@id='middle']/div/span[6]/table/tbody/tr[3]/td[1]/span[2]")
    raw_State = driver.find_element_by_xpath("//*[@id='middle']/div/span[6]/table/tbody/tr[3]/td[2]/span[2]")
    raw_Zip = driver.find_element_by_xpath("//*[@id='middle']/div/span[6]/table/tbody/tr[3]/td[3]/span[2]")
    # 21. Extract data from webdriver objects, and place into python variables.
    ownerShortName = raw_Owner_Short_Name.text.strip()
    address1 = raw_Address1.text.strip()
    city = raw_City.text.strip()
    state = raw_State.text.strip()
    zipCode = raw_Zip.text.strip()

    # 22. Print the data items to our output file.
    stringData = '"' + accountNumber + '","' + parcelNumber + '","' + actualParcelNumber + '","' + taxArea + '","' + legalClass + '","' + unitOfMeasure + '",' + parcelSize + ',"' + ownerShortName + '","' + address1 + '","' + city + '","' + state + '","' + zipCode + '"\n'
    # print stringData
    outfile.write(stringData)
    # 23. Print a status message to the user.
    iCntr += 1
    print "%s (%d/%d) captured." % (parcelNumber, iCntr, totItems)
    logfile.write(parcelNumber + '(' + str(iCntr) + '/' + str(totItems) + ') captured.\n')   
    # 24. Back to main search page
    driver.find_element_by_link_text('Account Search').click()

# 25. Cleanup
outfile.close

logfile.close

CODE WALK-THROUGH


1. Imports
These are the same as in the previous post, with the exception of importing the NoSuchElementException class. This class is leveraged down in step 11 to determine if our search for a parcel ID returned no results.

2. Create the webdriver object
Again, this is the same as the previous post. We interact with the webdriver object to do things, and capture data.

3. Setup Input and Output files
This is also the same as the previous post. We need to specify our input and output files (read parcel numbers from input, write web data to output).

4. Change the mode of the output file from write to append.
This is also the same with one exception. Instead of opening our file as "write" (w), we open as "append". This way, if we run into an error, we can restart the script, and it will append to the existing data file (the "write" mode overwrites all data in the file each time the file is opened - not what we want).

5. We create a simple counter variable that increments with each record that we process.

6. As with the previous post, we loop through each row in the input file to do some set of tasks.

7. As before, we load the county page. We set a time.sleep() here to ensure the page loads, but this is the last time we will used a time.sleep().

8. Make an element visible
This was a new problem that popped up with the Apache County page. The content on the page pushed the submit button down below the bottom of the window such that it was not visible. If an element is not visible, selenium can't work with it. Imagine trying to click a button that is not visible on the page - you can't do it, your only option is to scroll down to the button, then click it. We do the same here. The 'false' parameter to the scrollIntoView() function tells Firefox to scroll down only until the entire object is visible, then stop (as opposed to placing the submit button in the middle or top of the screen).

9. Wait for the page to load
The driver.implicitly_wait() function solves a very big problem for us: 'how to ensure we don't try to start reading data items before the page is fully loaded, yet not wait indefinitely?'. The driver.implicitly_wait(x) function will wait 'x' second for the page to load completely, then allow the script to continue to the next statement. If x seconds has passed, and the page still has not loaded, a timeout will occur, and the script will throw an exception. This wait applies to the remainder of the script (every new page selected has 'x' seconds to completely load or risk a timeout), so we no longer require time.wait() function calls.

10. Search for a parcel number, and bring up the Account Summary page
Something here has changed since the previous blog post. Our input file no longer contains *only* parcel IDs. It now contains Account Numbers, and Parcel Numbers. We leverage some python magic to grab both numbers for the current record. We'll see how the two are used further down. We tell selenium to wait for the parcel box item to become visible on the page, then we fill it with the parcelNumber we obtained from the input file, and click the submit button to submit the query to the web server.

11. Check for 'parcel not found' error
In testing, a string of text of class 'warning' will be displayed in the search results error if the parcel ID being sought was not found. By searching for that error, we can handle the exception by writing a message to the screen and new log file, then continuing on to the next line in the input data. If the exception is raised (the error was *not* found), then we simply pass to exit the try clause, and continue with trying to capture the data.

12. Handle multiple results
For this dataset, there may be multiple rows with the same parcelNumber, but each will have a different account number. This is where we leverage the account number (the first field) from the input file. In the results, we search for one containing the account number we were provided. If we get a hit, we click that row. If that fails (an exception would be thrown), we just grab the first row of class type 'clickable'.

13. Filter certain records
At this point, we have enough information visible on the screen to determine whether or not we want this record. Since the business rule I was given states. 'capture properties that consist of only land, and these have a type of '02.R'', we can drop in a simple if statement to check the value of the property type. If it's not 02.R, print a message to the user (and log the same to the logfile), skip all remaining instructions, and continue on with the next row in the input file. (Otherwise, continue with the script).

14./17./20. Capture data as webdriver objects
In these three steps, we scrape the page looking for data, and grab it using a webdriver object. The caveat here is that if we try to move on to the next page, the web elements we just captured will disappear.

15./18./21. Capture the data items as python variables. In these steps, we do some simple processing on the data (trimming excess whitespace, removing comma and dollar signs from currency values, etc.), then store the result in a python variable. That way, once we move to the next screen, if the webdriver objects disappear (and they will), we have the data we need captured with python for writing to the data file.

16./19. Move to the next page
In these steps, we move between pages of data. Since the links are simple textual links, we locate them by the text specified to show on the page, then click the link to jump to that page. I would point out that the page load wait instruction we entered in step 9. applies to these page jumps, also. Processing will not continue until the page loads. if 'x' seconds specified in the line in step 9 have passed, the script will throw an exception.

22. Output the data
Here, we create a single string (by hand) of each data item we captured from the various pages. Then, we send that string to the output file.

23. We now increment the counter variable, and print a message to the user that the data for the specific parcel id has been captured. We also write a line of the same to the log file.

24. By clicking the "Account Search" link, we go back to the initial search page, thus setting us up for the next data item in the input file.

25. As with the code in the previous blog post, we clean up our data files prior to exit.

UPDATED XPATH IDENTIFICATION


The previous post leveraged a tool that after an update to Firefox, stopped working. I went out and located another tool, FirePath. One of the advantages of FirePath is that it integrates directly into FireBug (which, if you followed the previous post, you would already be using). FirePath is simply another tab in FireBug. When inspecting elements in a web page, simply highlight the element you are interested in, and click the FirePath tab to get the XPath reference.

WRAP-UP

And that's it. We now have a much more robust, and flexible script to scrape from the web. Of course, the script is not finished. There are still a wide array of errors that could occur, and thus would require an exception handler. For example I did run across a one-in-4000+/- instance in which the script broke, presumably do to a page load timeout (re-running the script starting with the parcel ID that previously caused an error succeeded).

REFERENCES


Selenium Python API Guide